обновлено 2 дня назад
Senior Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes/AWS/Terraform): Building and operating reliable cloud infrastructure and platform tooling with an accent on Kubernetes, observability, infrastructure as code, and AI-native operational workflows. Focus on defining SLOs and error budgets, resolving systemic reliability issues, hardening infrastructure with Security, and leading incident response.
Location: Fully , with Europe prioritised for diversity and timezone requirements
Company
builds an HR platform with automation and AI capabilities for a globally distributed workforce.
What you will do
- Lead discovery and delivery of complex reliability and infrastructure solutions with a high degree of autonomy.
- Contribute to platform architecture, tooling, roadmap decisions, and technical initiatives.
- Define and operate SLOs, SLIs, error budgets, alerting, observability, and the platform's operational strategy.
- Identify systemic issues and create reusable fixes, runbooks, and solutions for cross-team requests.
- Build AI-native workflows, reusable prompts, tooling, agent-ready systems, and secure-by-default engineering guardrails.
- Mentor engineers, participate in hiring and RFC discussions, collaborate with Security, and join incident response and on-call rotations.
Requirements
- Professional experience in SRE, DevOps, or Platform Engineering.
- Hands-on experience operating and scaling production Kubernetes clusters, Docker, and related tooling.
- Experience managing AWS or similar cloud infrastructure and strong Terraform infrastructure-as-code skills.
- Knowledge of SLOs, SLIs, error budgets, alerting strategies, OpenTelemetry, Grafana, Prometheus, and observability practices.
- Experience with CI/CD and deployment automation, plus proficiency in Golang and Bash or scripting.
- Practical use of AI in infrastructure, operations, or development, including agentic workflows with observable results.
Nice to have
- Experience with Elixir, Node, Python, or another backend programming language.
- Experience running and configuring Linux systems outside cloud environments.
- Defensive and offensive security knowledge.
Culture & Benefits
- Fully , async-first work with flexible working hours.
- Work-from-anywhere benefit and flexible paid time off.
- Sixteen weeks of paid parental leave and mental health support services.
- Budgets for coworking spaces, learning, wellness, home office equipment, and IT equipment.
- Stock options and a focus on fair, location-adjusted compensation.
- Inclusive environment with employee resource groups and interview accommodations when needed.
Hiring process
- Recruiter interview followed by a hiring manager interview.
- Async infrastructure exercise requiring approximately 2–4 hours, followed by a team interview and Bar Raiser interview.
- Executive interview, offer, and background check.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Platform Engineer (SRE) (AI)
1 день назад
Senior Site Reliability Engineer (AI)
7 000 - 12 000$
Aleo Alliance
5 дней назад
DevOps Engineer (AI)
200 000₽
Chess.com
6 дней назад
Senior Site Reliability Engineer
4 дня назад
Site Reliability Engineer - AI Enablement
22 часа назад