Назад
Company hidden
1 день назад

Principal Site Reliability Engineer (AI)

200 000 - 250 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Site Reliability Engineer (Kubernetes/AI): Defining the long-term strategy and architecture for cloud and on-premise Kubernetes platforms across AWS, Google Cloud, and RKE2 with an accent on reliability, scalability, automation, and operational consistency. Focus on building Infrastructure as Code and GitOps workflows, establishing SLO/SLI and error budget frameworks, leading critical incidents, and applying AI-powered engineering capabilities to improve operational efficiency.

Location: Remote - US

Salary: $200,000–$250,000 USD annually, plus bonus, equity, and benefits.

Company

Publicly traded technology company operating in regulated sports betting and gaming.

What you will do

  • Define and execute the long-term Kubernetes platform strategy across Google Kubernetes Engine, Amazon Elastic Kubernetes Service, RKE2, and on-premise environments.
  • Lead architectural decisions covering cluster lifecycle management, networking, identity and access management, observability, autoscaling, capacity planning, and cost optimization.
  • Direct large-scale platform initiatives across multiple engineering teams and establish technical standards and measurable reliability outcomes.
  • Build automation-first infrastructure with Infrastructure as Code, GitOps, self-healing systems, and internal platform tooling.
  • Establish SLO, SLI, and error budget frameworks aligned with business priorities.
  • Lead critical incidents, drive post-incident improvements, mentor senior engineers, and promote responsible AI adoption in engineering workflows.

Requirements

  • Bachelor’s degree in Computer Science or a related technical field.
  • At least 8 years of experience designing, operating, and scaling distributed cloud and on-premise infrastructure, including 3 years at Staff, Principal, or equivalent technical leadership level.
  • Deep production expertise with Kubernetes, including architecture, networking, storage, security, operators, lifecycle management, and large-scale operations.
  • Extensive AWS and Google Cloud experience with Infrastructure as Code tools such as Terraform or Pulumi.
  • Strong software development skills in Go, Python, or both, plus GitOps, CI/CD, observability, distributed systems, Linux, and reliability engineering.
  • Exceptional communication and leadership skills, including mentoring engineers and influencing technical strategy. A gaming license issued by the appropriate state agency may be required.

Nice to have

  • Experience in regulated industries or hybrid cloud environments.
  • Open-source contributions or cloud certifications.

Culture & Benefits

  • Remote work within the United States.
  • Bonus, equity, and benefits applicable to the role.
  • Support with the gaming license process when required.
  • Focus on innovation, operational excellence, engineering productivity, and responsible AI adoption.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →