1 день назад
Site Reliability Engineer (Kubernetes)
123 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes): Building, operating, and scaling a multi-region SaaS platform across AWS, GCP, and Azure with an accent on Kubernetes infrastructure, service mesh architectures, and high-availability data layers. Focus on automating GitOps deployments, improving observability and incident response, and solving complex reliability, scalability, and resilience challenges in a 24/7 production environment.
Location: Remote, Washington, United States
Salary: $123,000–$150,000 per year; compensation varies by location, skills, experience, and role level. Benefits may vary by location.
Company
Develops API and AI connectivity technologies and a unified platform for securing, managing, governing, and monetizing API and AI model traffic.
What you will do
- Operate and scale a multi-region, multi-cloud SaaS platform across AWS, GCP, and Azure.
- Build Kubernetes infrastructure and deployment workflows with Terraform, Terragrunt, Helm, and ArgoCD.
- Design and optimize highly available, low-latency data and caching layers using PostgreSQL, Redis, ClickHouse, and Druid.
- Operate API gateway and service mesh environments supporting hybrid and distributed architectures.
- Develop CI/CD pipelines and GitOps workflows, and improve observability with Datadog, Prometheus, Grafana, and Thanos.
- Participate in a 24/7 on-call rotation, lead scaling initiatives, and improve incident response, postmortems, and operational playbooks.
Requirements
- Bachelor’s degree in Computer Science or equivalent practical experience.
- Experience managing enterprise-scale SaaS or PaaS systems in secure, multi-region, and multi-tenant environments.
- Deep Kubernetes expertise, including cluster and networking troubleshooting, fault tolerance, and scalability design.
- Strong proficiency with Infrastructure as Code, especially Terraform or Terragrunt, plus CI/CD and GitOps workflows.
- Programming experience with Go, Python, or Bash; solid knowledge of Linux/Unix, DNS, TLS/SSL, HTTP, load balancers, and distributed systems.
- Experience with API gateways, service meshes, Kafka, observability platforms, and 24/7/365 production support.
Nice to have
- Experience with Gateway, Mesh, ClickHouse, Druid, PostgreSQL, or Redis in multi-region environments.
- Knowledge of AWS networking, Azure VNet, or GCP NCC.
- Experience with disaster recovery, resiliency testing, and compliance-driven reliability practices.
Culture & Benefits
- Remote work supporting a global platform.
- US-based employees are typically offered healthcare benefits, a 401(k) plan, short- and long-term disability benefits, and life and AD&D insurance.
- Hands-on work with production SaaS systems serving thousands of customers across multiple regions and clouds.
- Operational work includes continuous improvement of reliability, security, performance, and cost efficiency.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
16 минут назад
Platform Engineer (AWS)
170 000 - 210 000$
5 часов назад
Senior DevOps Engineer (Kubernetes)
140 000 - 192 500CAD
7 дней назад
Senior Cloud Engineer (AWS/Kubernetes)
215 000 - 240 000$
12 часов назад
Site Reliability Engineer (AI Infrastructure)
200 000 - 240 000$
7 дней назад
Kubernetes Operator Engineer (AI)
12 часов назад
Senior Site Reliability Engineer (Cloud Infrastructure)
150 000 - 172 000$