1 день назад
Principal Site Reliability Engineer (AI)
200 000 - 250 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Site Reliability Engineer (Kubernetes/AI): Defining the long-term strategy and architecture for cloud and on-premise Kubernetes platforms across AWS, Google Cloud, and RKE2 with an accent on reliability, scalability, automation, and operational consistency. Focus on building Infrastructure as Code and GitOps workflows, establishing SLO/SLI and error budget frameworks, leading critical incidents, and applying AI-powered engineering capabilities to improve operational efficiency.
Location: Remote - US
Salary: $200,000–$250,000 USD annually, plus bonus, equity, and benefits.
Company
Publicly traded technology company operating in regulated sports betting and gaming.
What you will do
- Define and execute the long-term Kubernetes platform strategy across Google Kubernetes Engine, Amazon Elastic Kubernetes Service, RKE2, and on-premise environments.
- Lead architectural decisions covering cluster lifecycle management, networking, identity and access management, observability, autoscaling, capacity planning, and cost optimization.
- Direct large-scale platform initiatives across multiple engineering teams and establish technical standards and measurable reliability outcomes.
- Build automation-first infrastructure with Infrastructure as Code, GitOps, self-healing systems, and internal platform tooling.
- Establish SLO, SLI, and error budget frameworks aligned with business priorities.
- Lead critical incidents, drive post-incident improvements, mentor senior engineers, and promote responsible AI adoption in engineering workflows.
Requirements
- Bachelor’s degree in Computer Science or a related technical field.
- At least 8 years of experience designing, operating, and scaling distributed cloud and on-premise infrastructure, including 3 years at Staff, Principal, or equivalent technical leadership level.
- Deep production expertise with Kubernetes, including architecture, networking, storage, security, operators, lifecycle management, and large-scale operations.
- Extensive AWS and Google Cloud experience with Infrastructure as Code tools such as Terraform or Pulumi.
- Strong software development skills in Go, Python, or both, plus GitOps, CI/CD, observability, distributed systems, Linux, and reliability engineering.
- Exceptional communication and leadership skills, including mentoring engineers and influencing technical strategy. A gaming license issued by the appropriate state agency may be required.
Nice to have
- Experience in regulated industries or hybrid cloud environments.
- Open-source contributions or cloud certifications.
Culture & Benefits
- Remote work within the United States.
- Bonus, equity, and benefits applicable to the role.
- Support with the gaming license process when required.
- Focus on innovation, operational excellence, engineering productivity, and responsible AI adoption.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Site Reliability Engineer (Kubernetes)
123 000 - 150 000$
7 дней назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
4 дня назад
Senior Site Reliability Engineer (Kubernetes)
125 000 - 145 000$
1 день назад
Staff Site Reliability Engineer (AI/Blockchain)
195 000 - 257 500$
2 дня назад
Sr. Site Reliability Engineer (Healthcare)
125 000 - 145 000$
5 дней назад