обновлено 2 часа назад
Principal Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Site Reliability Engineer (Kubernetes/AI): Define the long-term strategy for cloud and on-premise infrastructure supporting a large-scale sports betting and gaming platform with an accent on Kubernetes architecture, reliability engineering, and automation. Focus on leading platform initiatives, building self-healing infrastructure with Infrastructure as Code and GitOps, improving incident response with AI-powered capabilities, and mentoring senior engineers.
Company
is a publicly traded technology company operating in regulated sports betting and gaming.
What you will do
- Define the long-term strategy for Kubernetes platforms across Google Kubernetes Engine, Amazon Elastic Kubernetes Service, RKE2, and on-premise environments.
- Drive architecture for cluster lifecycle management, networking, identity and access management, observability, autoscaling, capacity planning, and cost optimization.
- Lead large-scale platform initiatives across engineering teams and establish technical standards and measurable reliability outcomes.
- Develop service level objectives, service level indicators, error budget frameworks, Infrastructure as Code, GitOps workflows, self-healing systems, and internal platform tooling.
- Lead critical incidents, drive post-incident improvements, and strengthen resilience through automation and operational excellence.
- Mentor senior engineers, conduct architecture reviews, and influence technical strategy across the organization, including responsible adoption of AI-powered engineering capabilities.
Requirements
- Bachelor's degree in Computer Science or a related technical field.
- At least 8 years of experience designing, operating, and scaling distributed cloud and on-premise infrastructure, including at least 3 years at Staff, Principal, or equivalent technical leadership level.
- Deep production expertise with Kubernetes, including architecture, networking, storage, security, operators, lifecycle management, and large-scale operations.
- Extensive AWS and Google Cloud experience with Infrastructure as Code tools such as Terraform or Pulumi.
- Strong software development experience in Go, Python, or both, plus experience with GitOps, CI/CD, observability, distributed systems, Linux, and reliability engineering.
- Exceptional communication and leadership skills, including mentoring, cross-functional alignment, and technical strategy ownership.
Nice to have
- Experience in regulated industries or hybrid cloud environments.
- Contributions to open-source projects.
- Cloud certifications.
Culture & Benefits
- Technology-focused environment centered on innovation and emerging technologies.
- Opportunity to shape infrastructure strategy for a demanding sports betting and gaming platform.
- Support for obtaining a gaming license when required for the role.
- Equal employment opportunity and a commitment to a discrimination-free workplace.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
Site Reliability Engineer (Fintech)
87 400 - 123 400$
5 дней назад
Sr. Site Reliability Engineer (Kubernetes/AWS)
6 дней назад
Staff Site Reliability Engineer (GCP/Kubernetes)
4 часа назад
Staff Site Reliability Engineer (Kubernetes)
3 дня назад
Platform Site Reliability Engineer (SRE)
Wheely
5 дней назад
Site Reliability Engineer (AWS/Kubernetes)
5 000€