2 часа назад
Principal Site Reliability Engineer (Kubernetes)
200 000 - 250 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Site Reliability Engineer (Kubernetes): Shaping the long-term strategy for cloud and on-premise infrastructure supporting a large-scale sports betting and gaming platform with an accent on Kubernetes, reliability engineering, and infrastructure automation. Focus on defining platform architecture, leading cross-team initiatives, improving incident resilience, and adopting AI-powered engineering capabilities.
Location: Boston, Massachusetts, United States
Salary: $200,000–$250,000 base salary per year, plus bonus, equity, and benefits.
Company
is a publicly traded technology company operating in the regulated sports betting and gaming industry.
What you will do
- Define and execute the long-term Kubernetes platform strategy across Google Kubernetes Engine, Amazon Elastic Kubernetes Service, RKE2, and on-premise environments.
- Lead architectural decisions covering cluster lifecycle management, networking, identity and access management, observability, autoscaling, capacity planning, and cost optimization.
- Direct large-scale platform initiatives across multiple engineering teams, establishing technical standards and measurable reliability outcomes.
- Develop service level objectives, service level indicators, and error budget frameworks aligned with business priorities.
- Build automation-first infrastructure using Infrastructure as Code, GitOps, self-healing systems, and internal platform tooling.
- Lead critical incidents, improve platform resilience, adopt AI-powered engineering capabilities, and mentor senior engineers.
Requirements
- Bachelor's degree in Computer Science or a related technical field.
- At least 8 years of experience designing, operating, and scaling distributed cloud and on-premise infrastructure, including at least 3 years at Staff, Principal, or equivalent technical leadership level.
- Deep production expertise with Kubernetes, including architecture, networking, storage, security, operators, lifecycle management, and large-scale operations.
- Extensive AWS and Google Cloud Platform experience with Infrastructure as Code tools such as Terraform or Pulumi.
- Strong software development experience in Go, Python, or both, plus GitOps, CI/CD, observability, distributed systems, Linux, and reliability engineering.
- Exceptional communication and leadership skills, including mentoring, technical strategy, and cross-functional influence. A gaming license may be required as a condition of employment.
Nice to have
- Experience in regulated industries or hybrid cloud environments.
- Open-source contributions or cloud certifications.
Culture & Benefits
- Full-time employment with base salary, bonus, equity, and benefits.
- Work spans cloud and on-premise infrastructure supporting a regulated gaming platform.
- Opportunities to influence platform strategy, engineering standards, and AI adoption across the organization.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
Okta
2 дня назад
Staff Site Reliability Engineer (Splunk)
174 000 - 239 000$
Nscale
1 день назад
Senior Site Reliability Engineer (AI Infrastructure Operations)
170 000 - 265 000$
6 дней назад
Site Reliability Engineer (AWS)
180 000 - 220 000$
Okta
1 день назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
5 дней назад