27 дней назад
VP, Site Reliability Engineering (SRE & Observability Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
VP, Site Reliability Engineering (SRE & Observability Platform) (SRE, Observability, AI): Building a central SRE and observability platform organization with an accent on OpenTelemetry, SLOs, incident management, and agentic engineering. Focus on developing AI agents for instrumentation, incident investigation, on-call support, and guarded remediation while reducing production incidents across a large engineering estate.
Location: US FL JAX 347
Company
provides technology and services for financial services organizations.
What you will do
- Found and lead a central SRE platform organization covering SRE, platform engineering, AI engineering, enablement, and reliability champions.
- Own the observability and reliability platform, including OpenTelemetry telemetry pipelines, metrics, logs, traces, dashboards, alerting, SLOs, error budgets, incident management, and RCA.
- Build paved-road tooling with shared instrumentation SDKs, templates, and dashboards, alerts, and SLOs as code.
- Develop an agentic engineering stack for instrumentation, incident investigation, configuration generation, on-call assistance, and guarded remediation.
- Deliver AI agents to product engineers so they can instrument services, create SLOs and observability assets, investigate incidents, and support on-call operations.
- Own the roadmap, budget, vendor strategy, reliability governance, AI safety controls, and executive reporting.
Requirements
- 15+ years of software engineering experience, including 7+ years leading SRE, platform, infrastructure, or AI organizations at scale and managing managers.
- Proven success improving reliability and reducing production incidents across large, complex, multi-team environments.
- Production experience with AI or agentic systems, including LLMs, agent frameworks, orchestration, retrieval, evaluations, guardrails, and AI observability.
- Deep expertise in OpenTelemetry, metrics, logs, traces, SLOs, error budgets, incident management, Kubernetes, and infrastructure as code.
- Experience establishing AI governance with least-privilege access, human-in-the-loop controls, blast-radius limits, and evaluation-gated autonomy.
- Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
Nice to have
- Experience building internal AI or agent developer tooling adopted by engineering teams at scale.
- Experience in financial services or another regulated, high-availability, high-compliance environment.
- Familiarity with Prometheus, Grafana, observability SaaS, PagerDuty, SLO tooling, and LLM or agent observability platforms.
- Experience establishing error-budget policy and a blameless postmortem culture.
- Advanced degree in a relevant field.
Culture & Benefits
- Executive-sponsored reliability and AI transformation with committed funding.
- Reliability is treated as a product, with adoption driven through paved roads, enablement, office hours, and measurable results rather than mandates.
- Work includes building a blameless, high-trust engineering culture with psychological safety.
- Success is measured through reduced Sev1 and Sev2 incidents, lower mean time to resolution, SLO coverage, reduced alert noise, and healthier on-call operations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →