Назад
Company hidden
6 дней назад

Staff Engineer (SRE)

110 000 - 230 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Engineer (SRE/Incident Management): Building and operating automation, shared services, APIs, dashboards, and data pipelines that improve incident detection, response, recovery, and reliability across critical distributed platforms with an accent on observability, resilience, and production operations. Focus on leading high-severity incident response, performing root cause analysis, designing cross-team architecture, and driving 24x7 reliability improvements across Kubernetes and cloud environments.

Location: Bethesda, Maryland, United States

Annual salary: $110,000–$230,000

Company

hirify.global is a large United States auto insurer and a member of the Berkshire Hathaway family of companies, serving millions of customers nationwide.

What you will do

  • Design, develop, and operate automation, self-service tools, dashboards, and data pipelines for incident management, on-call, paging, and troubleshooting.
  • Build shared services, APIs, integrations, and data contracts that standardize incident response and reduce operational risk.
  • Lead technical response during high-severity incidents, including troubleshooting, impact analysis, cross-team coordination, and safe service restoration.
  • Lead post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements.
  • Drive architecture reviews, deployment safety, CI/CD, infrastructure as code, observability, testing, and production readiness across multiple teams.
  • Mentor engineers and influence technical direction, operational practices, and engineering culture.

Requirements

  • 8+ years of professional software engineering experience, including platform engineering, reliability engineering, backend engineering, distributed systems, or operational tooling.
  • 6+ years of experience with architecture, system design, reliability, scalability, and technical leadership for production systems.
  • Hands-on proficiency with multiple languages, including Go, Java, Python, and C#, plus Kubernetes and serverless technologies such as Knative.
  • Experience with Azure, AWS, or another cloud provider; SQL and NoSQL technologies; data pipelines; analytics; and operational dashboards.
  • Experience with OpenTelemetry, Grafana, Datadog, Splunk, Azure Monitor, and PagerDuty or comparable observability and incident management platforms.
  • Must participate in a 24x7 on-call rotation and support high-severity production incidents.

Nice to have

  • Experience with Spark, Trino, Superset, Power BI, and AI-assisted development tools such as Claude Code, Cursor, or GitHub Copilot.
  • Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.

Culture & Benefits

  • Personalized development programs, mentorship, and certification assistance.
  • Inclusive and collaborative culture focused on shared success and continuous improvement.
  • Competitive pay, benefits, and flexibility supporting employee well-being.
  • Engineering practices emphasize ownership, operational excellence, psychological safety, and learning from incidents.

Hiring process

  • Selection considers the role's scope, responsibilities, experience, education, training, work location, and business factors.
  • hirify.global will not sponsor a new applicant for employment authorization for this position.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →