Назад
Company hidden
11 часов назад

Site Reliability Engineer (Kubernetes)

140 000 - 170 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Kubernetes): Deploying and operating Blitzy's self-hosted AI software development platform inside a secure, customer-controlled cloud environment with an accent on Kubernetes operations, observability, and infrastructure reliability. Focus on hardening restricted deployments, designing monitoring and incident response, benchmarking compute-intensive AI workloads, and partnering directly with customer infrastructure and security teams.

Location: Remote within the U.S., with occasional travel for key customer workshops. U.S. citizenship is required for customer badging.

Salary: $140,000–$170,000 base salary annually, plus bonus and equity commensurate with experience.

Company

hirify.global is an AI software development platform that autonomously converts enterprise requirements into production-ready code.

What you will do

  • Deploy, operate, and maintain the self-hosted platform in a secure customer-controlled cloud environment.
  • Own Kubernetes deployments, releases, upgrades, capacity planning, and performance benchmarking for compute-intensive AI workloads.
  • Design observability systems covering logging, metrics, tracing, alerting, and incident response within the customer security boundary.
  • Partner with customer infrastructure, security, governance, and platform teams on provisioning, reviews, documentation, and escalations.
  • Handle sensitive customer data according to security requirements and promote security best practices.
  • Feed operational lessons into the product and infrastructure roadmap for future public sector deployments.

Requirements

  • U.S. citizenship and ability to complete the customer background and badging process.
  • 3+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • Strong Kubernetes and container orchestration experience, including deployments in customer-controlled or restricted environments.
  • Experience operating in isolated, restricted, or highly regulated environments such as defense, government, or financial services.
  • Hands-on infrastructure-as-code experience with Terraform, Pulumi, or equivalent, plus at least one major cloud platform.
  • Expertise in observability, incident management, on-call practices, and scripting with Python, Go, Bash, or similar tools.

Nice to have

  • Experience with government-accredited or similarly certified cloud environments.
  • Familiarity with sensitive data and security frameworks for regulated industries.
  • Experience supporting AI/ML workloads or their infrastructure.
  • Forward-deployed, residency, or embedded-engineer experience at an enterprise customer site.
  • Experience in a high-growth startup environment.

Culture & Benefits

  • Greenfield reliability work shaping the deployment playbook for government and defense customers.
  • Direct influence on architecture and close collaboration with engineering, infrastructure, security, and customer teams.
  • Fast-paced, innovation-focused environment with strong customer ownership.
  • Bonus and equity are provided in addition to the base salary.
  • Emphasis on sustainable performance through sleep, movement, and restorative activities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →