1 день назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (Kubernetes/GCP): Building and operating resilient, observable infrastructure for a software creation platform serving millions of developers with an accent on Kubernetes performance, distributed systems, and infrastructure automation. Focus on designing monitoring and SLO/SLI systems, leading incident response, debugging complex failures, and improving reliability through Python or Go tooling.
Location: Remote - United States
Salary: $250,000–$325,000 base annually, plus equity.
Company
Replit is an agentic software creation platform that enables people to build applications using natural language and serves millions of developers worldwide.
What you will do
- Architect and implement monitoring, logging, tracing, dashboards, and metrics for proactive visibility into system health.
- Define and manage Service Level Objectives and Service Level Indicators with product and engineering teams.
- Lead high-impact incident response, conduct blameless post-mortems, improve runbooks, and reduce Mean Time To Recovery.
- Build automation, CI/CD pipelines, infrastructure as code, and self-healing systems using tools such as Terraform and Pulumi.
- Optimize large-scale Kubernetes, Docker, and GCP deployments for performance, capacity, latency, and reliability.
- Review system designs, debug complex distributed-system problems, write internal tools in Python or Go, and mentor engineers across the company.
Requirements
- 8–10 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Infrastructure Engineering.
- Strong programming skills in Python or Go and experience writing high-quality, well-tested code.
- Deep experience designing, scaling, and maintaining distributed production services and service-oriented architectures.
- Expertise with Kubernetes, container orchestration, cloud-native technologies, and Google Cloud Platform.
- Proven experience with monitoring and observability platforms, incident management, infrastructure as code, and configuration management.
- Strong communication, mentoring, debugging, and technical leadership skills, including experience supporting engineers from junior to principal levels.
Nice to have
- Expertise with Prometheus, Grafana, Datadog, or OpenTelemetry.
- Significant experience with Go and Terraform.
- Experience building high-throughput, low-latency systems.
- Experience in rapid-growth startup environments.
- Experience writing company-facing blog posts and training materials.
Culture & Benefits
- Autonomous work environment with an emphasis on open and transparent communication.
- Competitive salary, equity, and a 401(k) program with a 4% match for US employees.
- Health, dental, vision, life, short-term disability, and long-term disability insurance.
- Paid parental, medical, and caregiver leave, plus flexible time off and holidays.
- Monthly wellness stipend, quarterly team gatherings, and commuter and office setup benefits where applicable.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
5 дней назад
Senior DevOps / Site Reliability Engineer (SRE) (Cybersecurity)
165 000 - 215 000$
8 дней назад
Senior Staff Site Reliability Engineer
232 338 - 290 422$
2 дня назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
4 дня назад
Sr. Site Reliability Engineer
160 000 - 180 000$
8 дней назад
Site Reliability Engineer (AWS)
120 000 - 185 000$