Назад
Company hidden
27 дней назад

VP, Site Reliability Engineering (SRE & Observability Platform)

Формат работы
onsite
Тип работы
fulltime
Грейд
head
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
VP, Site Reliability Engineering (SRE & Observability Platform) (SRE, Observability, AI): Building a central SRE and observability platform organization with an accent on OpenTelemetry, SLOs, incident management, and agentic engineering. Focus on developing AI agents for instrumentation, incident investigation, on-call support, and guarded remediation while reducing production incidents across a large engineering estate.

Location: US FL JAX 347

Company

hirify.global provides technology and services for financial services organizations.

What you will do

  • Found and lead a central SRE platform organization covering SRE, platform engineering, AI engineering, enablement, and reliability champions.
  • Own the observability and reliability platform, including OpenTelemetry telemetry pipelines, metrics, logs, traces, dashboards, alerting, SLOs, error budgets, incident management, and RCA.
  • Build paved-road tooling with shared instrumentation SDKs, templates, and dashboards, alerts, and SLOs as code.
  • Develop an agentic engineering stack for instrumentation, incident investigation, configuration generation, on-call assistance, and guarded remediation.
  • Deliver AI agents to product engineers so they can instrument services, create SLOs and observability assets, investigate incidents, and support on-call operations.
  • Own the roadmap, budget, vendor strategy, reliability governance, AI safety controls, and executive reporting.

Requirements

  • 15+ years of software engineering experience, including 7+ years leading SRE, platform, infrastructure, or AI organizations at scale and managing managers.
  • Proven success improving reliability and reducing production incidents across large, complex, multi-team environments.
  • Production experience with AI or agentic systems, including LLMs, agent frameworks, orchestration, retrieval, evaluations, guardrails, and AI observability.
  • Deep expertise in OpenTelemetry, metrics, logs, traces, SLOs, error budgets, incident management, Kubernetes, and infrastructure as code.
  • Experience establishing AI governance with least-privilege access, human-in-the-loop controls, blast-radius limits, and evaluation-gated autonomy.
  • Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.

Nice to have

  • Experience building internal AI or agent developer tooling adopted by engineering teams at scale.
  • Experience in financial services or another regulated, high-availability, high-compliance environment.
  • Familiarity with Prometheus, Grafana, observability SaaS, PagerDuty, SLO tooling, and LLM or agent observability platforms.
  • Experience establishing error-budget policy and a blameless postmortem culture.
  • Advanced degree in a relevant field.

Culture & Benefits

  • Executive-sponsored reliability and AI transformation with committed funding.
  • Reliability is treated as a product, with adoption driven through paved roads, enablement, office hours, and measurable results rather than mandates.
  • Work includes building a blameless, high-trust engineering culture with psychological safety.
  • Success is measured through reduced Sev1 and Sev2 incidents, lower mean time to resolution, SLO coverage, reduced alert noise, and healthier on-call operations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →