6 дней назад
Staff Engineer (SRE)
110 000 - 230 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Engineer (SRE/Incident Management): Building and operating automation, shared services, APIs, dashboards, and data pipelines that improve incident detection, response, recovery, and reliability across critical distributed platforms with an accent on observability, resilience, and production operations. Focus on leading high-severity incident response, performing root cause analysis, designing cross-team architecture, and driving 24x7 reliability improvements across Kubernetes and cloud environments.
Location: Bethesda, Maryland, United States
Annual salary: $110,000–$230,000
Company
is a large United States auto insurer and a member of the Berkshire Hathaway family of companies, serving millions of customers nationwide.
What you will do
- Design, develop, and operate automation, self-service tools, dashboards, and data pipelines for incident management, on-call, paging, and troubleshooting.
- Build shared services, APIs, integrations, and data contracts that standardize incident response and reduce operational risk.
- Lead technical response during high-severity incidents, including troubleshooting, impact analysis, cross-team coordination, and safe service restoration.
- Lead post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements.
- Drive architecture reviews, deployment safety, CI/CD, infrastructure as code, observability, testing, and production readiness across multiple teams.
- Mentor engineers and influence technical direction, operational practices, and engineering culture.
Requirements
- 8+ years of professional software engineering experience, including platform engineering, reliability engineering, backend engineering, distributed systems, or operational tooling.
- 6+ years of experience with architecture, system design, reliability, scalability, and technical leadership for production systems.
- Hands-on proficiency with multiple languages, including Go, Java, Python, and C#, plus Kubernetes and serverless technologies such as Knative.
- Experience with Azure, AWS, or another cloud provider; SQL and NoSQL technologies; data pipelines; analytics; and operational dashboards.
- Experience with OpenTelemetry, Grafana, Datadog, Splunk, Azure Monitor, and PagerDuty or comparable observability and incident management platforms.
- Must participate in a 24x7 on-call rotation and support high-severity production incidents.
Nice to have
- Experience with Spark, Trino, Superset, Power BI, and AI-assisted development tools such as Claude Code, Cursor, or GitHub Copilot.
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.
Culture & Benefits
- Personalized development programs, mentorship, and certification assistance.
- Inclusive and collaborative culture focused on shared success and continuous improvement.
- Competitive pay, benefits, and flexibility supporting employee well-being.
- Engineering practices emphasize ownership, operational excellence, psychological safety, and learning from incidents.
Hiring process
- Selection considers the role's scope, responsibilities, experience, education, training, work location, and business factors.
- will not sponsor a new applicant for employment authorization for this position.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Replit
10 дней назад
Engineering Manager (SRE)
250 000 - 325 000$
13 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
10 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
12 дней назад
Staff Engineer (SRE)
95 800 - 185 000$
8 дней назад
Staff Site Reliability Engineer (AI)
220 000 - 260 000$