4 дня назад
Sr. Staff Lead Site Reliability Engineer (AWS)
220 000 - 330 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. Staff Lead Site Reliability Engineer (AWS): Establishing and maturing SRE practices for cloud infrastructure and platform services with an accent on reliability targets, observability, incident response, and resilient operations. Focus on diagnosing complex distributed-system failures, building operational automation, and leading multi-quarter reliability roadmaps across engineering teams.
Location: San Mateo, California, United States; Workplace: On-site
Salary: $220,000–$330,000 per year, plus bonus, benefits, and equity for regular employees.
Company
is a venture-backed defense-tech company developing intelligent systems, including Hivemind autonomy software, autonomous aircraft, and simulation technologies.
What you will do
- Establish and mature SRE practices across cloud infrastructure and platform services.
- Define and implement SLIs, SLOs, reliability targets, monitoring, alerting, logging, and tracing.
- Lead technical incident response, root-cause analysis, and elimination of recurring failure modes.
- Improve resilience through automation, testing, capacity planning, and recovery engineering.
- Develop operational tooling that reduces manual work and partner with product and platform teams on reliability requirements.
- Define the short- and long-term SRE roadmap, distribute work, and mentor engineers across SRE and Cloud Engineering.
Requirements
- 7+ years of experience in SRE, software engineering, infrastructure engineering, or a related field.
- Experience operating production services with availability and reliability requirements.
- Experience with SLIs, SLOs, monitoring, alerting, incident response, and root-cause analysis.
- Experience designing and operating infrastructure in AWS or another major cloud environment, including infrastructure-as-code and automated provisioning.
- Experience supporting containerized applications and distributed systems.
- Experience developing operational tooling with Python, Go, or a similar language, and leading technical vision over multi-quarter timelines.
Nice to have
- Experience establishing or maturing an SRE function.
- Kubernetes and cloud-native observability experience.
- Experience with regulated or compliance-driven environments.
- Capacity planning, performance analysis, cloud cost management, and shared infrastructure experience.
- Experience building roadmaps in ticketing systems and mentoring engineers.
Culture & Benefits
- Full-time regular employees receive benefits, bonus eligibility, and equity.
- Employment offers are subject to a cleared background and possible reference check.
- Artificial intelligence tools may support parts of the hiring process, while final decisions are made by humans.
- is an equal opportunity and affirmative action employer.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Reddit
5 дней назад
Staff Site Reliability Engineer, Ads
217 000 - 303 900$
2 дня назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
4 дня назад
Site Reliability Engineer (AWS)
180 000 - 220 000$
Anthropic
2 дня назад
Staff+ Site Reliability Engineer (Safeguards ML Infra)
320 000 - 485 000$
2 дня назад
Member of Technical Staff, Infrastructure Engineer (AI)
175 000 - 240 000$
6 дней назад