Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr Platform Monitoring Engineer (Cloud Observability): Building customer-focused monitoring solutions, alerting pipelines, and observability workflows for the Databricks Platform with an accent on incident detection, reliability, and cross-functional response. Focus on investigating complex production incidents, analyzing root causes across infrastructure and cloud providers, and automating monitoring patterns that reduce customer impact.
Location: United States
Company
Databricks builds and operates a data and AI infrastructure platform used by organizations worldwide to develop and scale data, AI, analytics, and agent applications.
What you will do
- Lead platform incident investigations and coordinate cross-functional teams through detection, mitigation, and resolution.
- Conduct root cause analyses across infrastructure, services, and cloud providers, identifying systemic patterns and prevention opportunities.
- Design and implement customer-focused alerting pipelines and end-to-end observability workflows.
- Build automation tools, reusable monitoring patterns, and solutions for platform reliability gaps.
- Mentor junior engineers on observability patterns, alert design, and service health metrics.
- Participate in the on-call rotation.
Requirements
- At least 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or in a similar role.
- Production experience with AWS, Azure, or GCP, plus Docker and Kubernetes.
- Hands-on experience with monitoring, logging, and alerting tools such as ELK, Prometheus, Grafana, or PagerDuty.
- Ability to architect solutions that correlate metrics, logs, and traces.
- Strong Python or similar-language skills for building production-quality automation tools.
- BS, Master's, or PhD in Computer Science, Computer Engineering, or a related engineering field.
Culture & Benefits
- Customer-obsessed engineering culture focused on solving complex technical challenges.
- Work on a large-scale data and AI infrastructure platform.
- Comprehensive employee benefits and perks, with details varying by region.
- Commitment to diversity, inclusion, and equal employment opportunity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Sr. Staff Observability Software Engineer (Kubernetes)
17 400 - 299 000$
9 часов назад
Senior Site Reliability Engineer (Observability)
5 дней назад
Site Reliability Engineer (AI)
Nscale
6 дней назад
Senior Observability Platform Engineer (AI)
160 000 - 230 000$
14 часов назад
Principal Cloud Platform Engineer (AI)
144 000 - 189 000$
Cloud.ru
7 дней назад