4 дня назад
Sr. Staff Lead Site Reliability Engineer (R5803)
183 000 - 275 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. Staff Lead Site Reliability Engineer (SRE/cloud infrastructure): Establishing and maturing reliability practices across cloud infrastructure and platform services with an accent on observability, incident response, resilience, and operational automation. Focus on defining SLIs and SLOs, diagnosing complex distributed-system failures, leading root-cause analysis, and building tooling that improves recovery and reduces manual operations.
Location: San Diego, California; on-site
Salary: $183,000–$275,000 per year, plus bonus, benefits, and equity for regular full-time employees.
Company
is a venture-backed defense-tech company developing intelligent systems, including Hivemind autonomy software, V-BAT and X-BAT aircraft, and simulation and synthetic reality technologies.
What you will do
- Establish and mature the SRE function across cloud infrastructure and platform services.
- Define and implement SLIs, SLOs, monitoring, alerting, logging, and tracing.
- Lead technical incident response, complex failure investigation, and root-cause analysis.
- Improve system resilience through automation, testing, capacity planning, and recovery engineering.
- Partner with product and platform teams to incorporate reliability requirements into system design.
- Define the short- and long-term SRE roadmap, distribute work, and mentor engineers.
Requirements
- 7+ years of experience in SRE, software engineering, infrastructure engineering, or a related field.
- Experience operating production services with defined availability and reliability requirements.
- Experience with AWS or another major cloud environment, infrastructure as code, and automated provisioning.
- Experience supporting containerized applications and distributed systems.
- Operational tooling or automation experience using Python, Go, or a similar language.
- Experience leading incident response, root-cause analysis, and technical execution across multi-quarter timelines.
Nice to have
- Experience establishing or maturing an SRE function.
- Experience with Kubernetes and cloud-native observability systems.
- Experience in regulated or compliance-driven environments.
- Experience with capacity planning, performance analysis, cloud cost management, or shared infrastructure.
- Experience building roadmaps in ticketing systems and mentoring engineers.
Culture & Benefits
- Regular full-time employees receive benefits, bonus eligibility, and equity.
- Compensation is determined by skills, experience, certifications, licenses, and work location.
- Offers are contingent on a cleared background and possible reference check.
- is an equal opportunity and affirmative action employer.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Reddit
5 дней назад
Staff Site Reliability Engineer, Ads
217 000 - 303 900$
2 дня назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
6 дней назад
Infrastructure Engineer (AI)
200 000 - 300 000$
Anthropic
2 дня назад
Staff+ Site Reliability Engineer (Safeguards ML Infra)
320 000 - 485 000$
2 дня назад
Member of Technical Staff, Infrastructure Engineer (AI)
175 000 - 240 000$
6 дней назад