Назад
4 дня назад

Senior Software Engineer - Incident Insights & Readiness (SRE)

192 000 - 240 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer - Incident Insights & Readiness (SRE) (Go/Python/Kubernetes): Building software, tooling, and operational frameworks that improve on-call rotations, incident response, and post-mortem learning across Datadog with an accent on distributed systems, reliability engineering, and blameless incident management. Focus on designing incident platforms, analyzing systemic risks, leading cross-functional reliability initiatives, and coaching engineers through complex failure modes.

Location: Hybrid in Boston, Massachusetts, USA or New York, New York, USA

Salary: $192,000–$240,000 USD per year, plus equity and variable compensation

Company

Datadog provides an observability and security platform for managing applications, infrastructure, data, models, and security at scale.

What you will do

  • Own and improve the company-wide on-call experience, including best practices, rotation support, and compensation platforms.
  • Design and implement software that streamlines incident response and collaborate with product teams to improve operational processes.
  • Support post-mortems and incident reviews, identify systemic risks, and turn operational insights into reliability improvements.
  • Facilitate blameless learning across engineering teams and help share incident-related improvements throughout the organization.
  • Provide technical leadership, mentoring, design reviews, coaching, and operational excellence guidance.
  • Lead cross-functional initiatives that improve reliability and operational excellence across Datadog.

Requirements

  • At least 5 years of experience building software that solves real user problems.
  • Professional experience with Go and Python; familiarity with TypeScript, distributed systems, Kubernetes, and complex failure modes.
  • Experience owning ambiguous technical problems from design through delivery while balancing quality and pragmatic execution.
  • Experience analyzing incidents, participating in on-call rotations, and improving incident response processes.
  • Experience mentoring engineers, leading cross-functional initiatives, and influencing technical direction without formal authority.
  • Strong English communication skills are required for collaboration across teams.

Nice to have

  • Experience serving as an incident commander or incident coordinator.
  • Background in software engineering, site reliability engineering, production engineering, infrastructure, or related reliability-focused roles.

Culture & Benefits

  • Hybrid workplace emphasizing office collaboration and work-life harmony.
  • New-hire stock equity, an employee stock purchase plan, and competitive benefits.
  • Professional development, product training, career pathing, mentoring, and buddy programs.
  • Inclusive employee resource groups, community guilds, and internal inclusion discussions.
  • Global mental health benefits, healthcare, dental coverage, parental planning, paid time off, fitness reimbursements, and a 401(k) plan with matching.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →