7 часов назад
Engineering Manager, SRE (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Engineering Manager, SRE (AI): Leading a distributed Site Reliability Engineering team responsible for Kubernetes, AWS, PostgreSQL, CI infrastructure, observability, and reliability practices with an accent on people leadership, AI infrastructure, and operational excellence. Focus on scaling SLOs and error budgets, balancing on-call operations with delivery, and building reliable infrastructure and incident-response practices across a global organization.
Location: Anywhere in the world; coverage is strongest in EMEA and APAC, with opportunities to support the Americas.
Company
builds a global HR platform with automation and AI capabilities.
What you will do
- Lead the Site Reliability Engineering team with responsibility for career development, performance, hiring, team health, and progression.
- Set SRE goals, prioritize delivery, and manage the support rotation and on-call model.
- Own core infrastructure across Kubernetes, AWS, PostgreSQL, DNS, TLS, and CI systems.
- Advance SLOs, error budgets, incident response, observability, and reliability practices.
- Partner with Security on threat management, patching, infrastructure controls, audits, and compliance.
- Represent SRE across engineering and senior leadership while managing vendor relationships.
Requirements
- Experience leading an SRE, infrastructure, or platform engineering team and developing direct reports.
- Hands-on site reliability, DevOps, or cloud infrastructure background with Kubernetes in production and AWS at meaningful scale.
- Experience building, enabling, or scaling AI infrastructure.
- Knowledge of observability, Terraform, CI/CD systems such as GitLab CI, GitHub Actions, or Jenkins, Docker, and shell scripting.
- Experience with incident response, on-call operations, SLOs, error budgets, and regulated environments.
- Strong prioritization, written communication, conflict resolution, and cross-team collaboration skills in a fully distributed async organization.
Nice to have
- Backend development experience with Elixir, Java, Clojure, Node, Python, or a similar language.
- Experience with OpenTelemetry, distributed tracing, Honeycomb, PostgreSQL or Aurora operations, Linux, security, FinOps, or team growth.
Culture & Benefits
- Work from anywhere with flexible working hours in an async environment.
- Flexible paid time off and 16 weeks of paid parental leave.
- Budgets for co-working spaces, learning, wellness, home office equipment, and IT equipment.
- Mental health support services and stock options.
- Emphasis on life-work balance, autonomy, inclusion, and employee growth.
Hiring process
- Recruiter interview followed by a hiring manager interview.
- Scenario interview with a peer team leader, executive interview, and Bar Raiser interview.
- Offer stage includes prior employment verification.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 часов назад
Manager, Software Engineering (SRE)
11 часов назад
Senior Platform Engineering Manager (AI)
6 дней назад
Senior Manager, Software Engineering (AI)
200 000 - 235 000$
Okta
2 дня назад
Manager, Site Reliability Engineering (AWS/Kubernetes)
204 000 - 306 000$
Okta
1 день назад
Manager, Site Reliability Engineering (AWS/Kubernetes)
204 000 - 306 000$
6 дней назад