3 дня назад
Staff Site Reliability Engineer (GCP/Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (GCP/Kubernetes): Designing and operating the platform behind Finalsite's Composer CMS in a GCP-primary, multi-cloud environment with an accent on cloud architecture, Kubernetes reliability, infrastructure as code, observability, and network security. Focus on building highly available systems, leading disaster recovery and high-severity incident response, optimizing cloud costs, and mentoring senior engineers.
Location: 100% remote within the United States; current and continued U.S. residency required
Company
provides an integrated platform for K–12 schools, including websites, communications, mobile apps, enrollment, and marketing services.
What you will do
- Set technical direction for the Composer CMS platform and platform-wide standards in a GCP-primary, multi-cloud environment.
- Guide cloud architecture, scalability, capacity planning, cost efficiency, and highly available system design.
- Own the GKE Kubernetes platform, container reliability, network architecture, global routing, connectivity, and Cloudflare edge strategy.
- Drive reusable Terraform and Terragrunt infrastructure-as-code modules and developer enablement patterns.
- Build observability practices, define monitoring and alerting standards, and establish SLOs.
- Lead disaster recovery, high-severity incident response, on-call participation, change management, and post-incident improvements while mentoring senior engineers.
Requirements
- Expert-level GCP and broad infrastructure knowledge across the full stack.
- Expertise in Terraform, Terragrunt, Kubernetes, containerization, and Helm.
- Experience designing and operating CI/CD pipelines with GitLab or similar tools.
- Strong network architecture, routing, cross-cloud connectivity, edge security, and traffic management experience, including Cloudflare or similar solutions.
- Deep observability knowledge, incident leadership experience, and experience architecting highly available, fault-tolerant systems.
- Staff-level experience, including mentoring senior engineers, plus the ability to read and write basic code in environments such as Ruby/Rails, Java, and Python.
Nice to have
- Experience with cloud cost allocation and control.
- Capacity planning and infrastructure forecasting experience at scale.
- Familiarity with AWS and/or Azure; the environment is GCP-primary.
Culture & Benefits
- Fully remote employment for eligible U.S. residents.
- Collaborative, supportive culture centered on partnership and purpose.
- Competitive benefits and professional development opportunities.
- Peer review, staged rollouts, and go/no-go gates support controlled change management.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Staff Site Reliability Engineer (Kubernetes)
7 дней назад
Staff Site Reliability Engineer (GCP/Kubernetes)
112 500 - 187 500$
5 дней назад
Lead Site Reliability Engineer (Kubernetes)
3 дня назад
Senior Site Reliability Engineer (Cloud-Native Infrastructure)
142 800 - 178 500$
2 дня назад
Cloud Engineer (AWS)
7 дней назад
Senior Manager of Site Reliability Engineering (AWS)
155 000 - 170 000$