1 день назад
Senior Engineer (SRE/Incident Management)
100 000 - 215 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Engineer (SRE/Incident Management): Building and operating enterprise-critical platforms, automation, APIs, dashboards, and data pipelines for incident management and service reliability with an accent on observability, distributed systems, and production operations. Focus on leading high-severity incident response, improving root cause analysis and post-incident actions, and designing resilient tooling for 24x7 mission-critical environments.
Location: Bethesda, Maryland, United States
Annual salary: $100,000–$215,000
Company
is a large United States auto insurer and a member of the Berkshire Hathaway family of companies.
What you will do
- Design, develop, and operate automation, self-service tools, dashboards, APIs, integrations, and data pipelines for incident management and on-call operations.
- Build reliable shared services and apply engineering standards across deployment, testing, observability, security, and production readiness.
- Lead technical response during high-severity incidents, coordinating troubleshooting, impact analysis, and safe service recovery.
- Conduct post-incident reviews, root cause analysis, corrective action planning, and reliability improvements.
- Develop runbooks, readiness criteria, triage models, and resilience practices for distributed systems.
- Partner with SRE, platform, product, infrastructure, security, and business stakeholders while mentoring engineers.
Requirements
- 4+ years of professional software engineering experience and experience building reliable production systems.
- Hands-on proficiency in multiple languages, including Go, Java, Python, or C#, with Kubernetes and serverless technologies such as KNative.
- Experience with SQL and NoSQL databases, cloud-native services, data pipelines, analytics, and operational dashboards.
- Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, or Azure Monitor, plus incident management platforms such as PagerDuty.
- Experience with Azure, AWS, GCP, or complex hybrid environments, and ownership of production systems operating 24x7.
- Must participate in a 24x7 on-call rotation and support high-severity production incidents.
Nice to have
- Experience with Spark, Trino, Superset, Power BI, or AI-assisted development tools such as Claude Code, Cursor, and GitHub Copilot.
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.
Culture & Benefits
- Personalized development programs, mentorship, and certification assistance.
- Inclusive and collaborative culture focused on shared success and continuous improvement.
- Competitive pay, benefits, and flexibility supporting employee well-being.
- Active ownership of systems through a “You Build It, You Run It” engineering model.
- will not sponsor a new applicant for employment authorization for this position.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
21 час назад
Sr Software Development Engineer, SRE (US Federal)
163 800 - 245 800$
1 день назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
7 дней назад
Site Reliability Engineering (SRE) Manager
106 000 - 130 600$
5 дней назад
DevOps & SRE Engineer (Kubernetes)
100 000 - 150 000$
1 день назад
Site Reliability Engineer Technical Lead (Kubernetes)
100 000 - 150 000$