5 дней назад
Senior Observability Engineer (Datadog)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Observability Engineer (Datadog): Building and operating an observability control plane for financial infrastructure with an accent on automated monitoring standards, service coverage, and incident-driven detection. Focus on auditing Ruby, Go, and Java services, reducing false-severity pages, and closing detection gaps through shared incident response and automation.
Location: Based in Lehi, Utah, United States; hybrid office collaboration is encouraged
Company
is a fintech company building technology that helps banks, credit unions, and fintechs deliver financial experiences to millions of people.
What you will do
- Build and operate an observability control plane using the Datadog API and Terraform to automate monitors, dashboards, tagging, and service onboarding.
- Define and audit observability standards for Ruby, Go, and Java services, including tags, golden signals, alert quality, and dashboard contracts.
- Produce detection and dashboard gap packs after significant incidents, then implement monitors and dashboards to close identified gaps.
- Track observability maturity, service-catalog health, ownership, SLO coverage, dashboard freshness, and monitoring trends through monthly reports.
- Tune alerting toward fewer false SEV1/2 pages and more actionable SEV3/4 alerts while coaching teams on Datadog cost and cardinality.
- Participate in the shared incident response and observability on-call rotation, investigating incidents and serving as Incident Commander or a supporting technical lead.
Requirements
- BS in Computer Science or equivalent experience.
- 5+ years of production observability, SRE, or DevOps experience.
- 5+ years of automation-first engineering with Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency.
- Experience debugging distributed systems across microservices, Kubernetes, and bare metal, including latency, connection pools, queues, cascading failures, NATS, RabbitMQ, PostgreSQL, and Redis.
- AI- and workflow-literate, with experience using or building scripted and AI-assisted workflows for reviews, audits, or documentation.
- Experience with shared on-call responsibilities and Incident Commander duties.
Nice to have
- Fintech experience and familiarity with Datadog; strong Grafana, Prometheus, Splunk, or New Relic experience is also considered.
- Knowledge of Google SRE practices, toil elimination, self-healing automation, and OpenTelemetry instrumentation.
- Experience influencing teams without direct authority and producing governance or compliance reports.
- Experience with incident.io, PagerDuty, or OpsGenie, plus Golang and Ruby on Rails.
Culture & Benefits
- Work in a high-performance environment focused on trust, accountability, and measurable results.
- Use the Utah office for collaboration, project kickoffs, and cross-functional strategy sessions.
- Lehi office perks include company-paid meals, a sports simulator, gym, mother’s lounge, and meditation room.
- Work on infrastructure supporting financial applications used by millions of people and billions of transactions.
- Inclusive workplace committed to equal opportunity and reasonable accommodations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (Kubernetes)
128 500 - 190 000$
Nscale
4 дня назад
Senior Observability Platform Engineer (AI)
160 000 - 230 000$
4 дня назад
Senior DevOps Engineer (AWS)
5 дней назад
Senior Platform/SRE Engineer (AWS)
5 дней назад
DevOps Engineer (Defense Technology)
125 000 - 160 000$
5 дней назад
Staff Site Reliability Engineer (AWS GovCloud)
158 500 - 230 000$