Назад
5 часов назад

Sr Platform Monitoring Engineer

Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr Platform Monitoring Engineer (Cloud Observability): Building customer-focused monitoring solutions, alerting pipelines, and observability workflows for the Databricks Platform with an accent on incident detection, reliability, and cross-functional response. Focus on investigating complex production incidents, analyzing root causes across infrastructure and cloud providers, and automating monitoring patterns that reduce customer impact.

Location: United States

Company

Databricks builds and operates a data and AI infrastructure platform used by organizations worldwide to develop and scale data, AI, analytics, and agent applications.

What you will do

  • Lead platform incident investigations and coordinate cross-functional teams through detection, mitigation, and resolution.
  • Conduct root cause analyses across infrastructure, services, and cloud providers, identifying systemic patterns and prevention opportunities.
  • Design and implement customer-focused alerting pipelines and end-to-end observability workflows.
  • Build automation tools, reusable monitoring patterns, and solutions for platform reliability gaps.
  • Mentor junior engineers on observability patterns, alert design, and service health metrics.
  • Participate in the on-call rotation.

Requirements

  • At least 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or in a similar role.
  • Production experience with AWS, Azure, or GCP, plus Docker and Kubernetes.
  • Hands-on experience with monitoring, logging, and alerting tools such as ELK, Prometheus, Grafana, or PagerDuty.
  • Ability to architect solutions that correlate metrics, logs, and traces.
  • Strong Python or similar-language skills for building production-quality automation tools.
  • BS, Master's, or PhD in Computer Science, Computer Engineering, or a related engineering field.

Culture & Benefits

  • Customer-obsessed engineering culture focused on solving complex technical challenges.
  • Work on a large-scale data and AI infrastructure platform.
  • Comprehensive employee benefits and perks, with details varying by region.
  • Commitment to diversity, inclusion, and equal employment opportunity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →