Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Engineering Manager (SRE/AI): Leading observability, incident management, load testing, performance engineering, and rollout infrastructure for AI-powered software creation systems with an accent on production reliability, measurable performance, and safe changes. Focus on building SRE platforms, diagnosing distributed-system bottlenecks, validating recovery under realistic load, and growing a high-ownership engineering team.
Location: Foster City, California, United States; hybrid
Salary: $250,000–$325,000 base salary plus equity
Company
Replit is an agentic software creation platform that enables people to build applications with natural language and AI.
What you will do
- Lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure.
- Build and operate metrics, logs, traces, alerting, SLOs, incident tooling, and reliability practices.
- Develop load and failure testing capabilities to validate critical paths, recovery, headroom, and production readiness.
- Lead performance investigations using profiling, telemetry, and load tests to identify and resolve bottlenecks with service owners.
- Review designs and production changes, debug difficult failure modes, and use AI coding tools while maintaining rigorous safeguards.
- Coach, hire, and develop engineers while measuring rollout safety, recovery time, latency, throughput, and test coverage.
Requirements
- Demonstrated experience managing, developing, hiring, and performance-managing engineers.
- Depth in distributed systems or reliability platforms, including deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.
- Experience leading consequential migrations or incidents and validating reliability and performance fixes under realistic conditions.
- Ability to build platform capabilities adopted by other teams and balance reliability, performance, engineering effort, and cost.
- Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.
- Experience with cloud cost attribution, capacity planning, provider coordination, and GCP.
Nice to have
- Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.
- Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.
Culture & Benefits
- Autonomous work environment with quarterly team gatherings.
- Competitive salary, equity, and a 401(k) program with a 4% match for US employees.
- Health, dental, vision, life, disability, parental, medical, and caregiver benefits.
- Flexible time off and holidays.
- Monthly wellness stipend, commuter benefits, office setup reimbursement, and office amenities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
7 дней назад
Staff Engineer (SRE)
95 800 - 185 000$
8 дней назад
Site Reliability Engineering (SRE) Manager (Azure)
139 700 - 232 900$
8 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
5 дней назад
Staff Engineer, Software Engineering (SRE Availability and Incident Management)
100 000 - 230 000$
8 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$