6 дней назад
Senior Staff Engineer - SRE (Incident Prevention)
110 000 - 260 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Staff Engineer - SRE (Incident Prevention): Building automation, self-service tools, dashboards, and data pipelines that improve correction-of-error workflows and reliability across distributed production systems with an accent on incident forensics, observability, and root cause analysis. Focus on identifying repeat incident patterns, leading high-severity incident response, and driving architecture and engineering improvements across teams.
Location: Bethesda, Maryland, United States
Annual salary: $110,000–$260,000
Company
is a large U.S. auto insurer and a member of the Berkshire Hathaway family of companies, serving millions of customers through technology and insurance services.
What you will do
- Design, develop, and operate automation, self-service tools, dashboards, and data pipelines that scale correction-of-error workflows.
- Lead and improve the company-wide correction-of-error and post-incident review process.
- Moderate weekly reviews of qualified high-severity incidents and help teams produce accurate, actionable analyses.
- Perform root cause analysis across distributed systems using logs, metrics, traces, and observability data.
- Identify reliability gaps involving alerts, monitoring, runbooks, testing, and deployment controls.
- Partner with application engineering, platform, SRE, and operations teams to improve technology strategies and roadmaps.
Requirements
- 10+ years of professional software engineering experience, including significant experience with architecture, system reliability, scalability, and technical leadership.
- Hands-on proficiency with multiple programming languages, including Go, Java, Python, and C#.
- Experience building production systems on Kubernetes and serverless technologies such as KNative in Azure and AWS.
- Experience with SQL and NoSQL databases, data pipelines, analytics, and operational dashboards.
- Experience with OpenTelemetry, Grafana, Datadog, Splunk, Azure Monitor, and PagerDuty.
- Must participate in a 24x7 on-call rotation and support high-severity production incidents.
Nice to have
- Experience with Spark, Trino, Superset, or Power BI.
- Experience with AI-assisted development tools such as Claude Code, Cursor, or GitHub Copilot.
- Experience with open-source frameworks, modern engineering practices, or platform technologies.
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.
Culture & Benefits
- Engineering culture focused on reliability, continuous improvement, and psychological safety.
- Senior technologists take part in high-severity incident leadership and production support.
- Personalized development programs, mentorship, and certification assistance.
- Inclusive and collaborative work environment with competitive benefits and flexibility.
- will not sponsor a new applicant for employment authorization for this position.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
6 дней назад
Sr Software Development Engineer, SRE (US Federal)
163 800 - 245 800$
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
12 дней назад
Site Reliability Engineering (SRE) Manager
106 000 - 130 600$
13 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
12 дней назад
Staff Engineer (SRE)
95 800 - 185 000$