Назад
Company hidden
19 часов назад

Staff Production Engineer (AI)

209 000 - 253 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Production Engineer (AI): Improving the reliability, scalability, and performance of an energy-efficient GPU cloud platform with an accent on SLOs, incident response, observability, and distributed systems. Focus on designing self-healing automation, strengthening disaster recovery, and solving complex reliability challenges across AI and HPC infrastructure.

Location: Sunnyvale, California, United States; on-site

Salary: $209,000–$253,000 annually plus bonus and restricted stock units.

Company

hirify.global builds vertically integrated, energy-efficient AI infrastructure and operates an AI-optimized cloud platform for demanding compute workloads.

What you will do

  • Define, measure, and improve availability metrics, service level indicators, and service level objectives for the cloud platform.
  • Lead production incident response, service disruption resolution, post-incident reviews, and root cause analysis.
  • Architect and improve observability using Prometheus, Grafana, Alertmanager, and OpenTelemetry.
  • Identify reliability risks and performance bottlenecks across distributed systems and GPU infrastructure.
  • Design automation, remediation tooling, and self-healing infrastructure to reduce operational toil and improve recovery times.
  • Partner with compute, networking, storage, and platform teams while mentoring engineers and promoting reliability practices.

Requirements

  • 8+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations.
  • Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems.
  • Strong Linux/Unix debugging skills across kernel and user space.
  • Knowledge of Kubernetes, distributed systems, virtualization, and cloud platforms such as AWS or GCP.
  • Experience with incident management, reliability frameworks, monitoring and observability, and infrastructure-as-code tools such as Terraform or Ansible.
  • Proficiency in Go, Python, C, or C++, with strong communication and cross-functional collaboration skills.

Nice to have

  • Experience leading Kubernetes or container orchestration platforms at scale.
  • Experience with operational readiness reviews, change management, structured root cause analysis, or automated remediation.
  • Interest in scaling AI or HPC infrastructure and mentoring Production Engineering teams.

Culture & Benefits

  • Health, dental, and vision insurance options, including HDHP and PPO plans.
  • Employer HSA contributions, paid parental leave, life insurance, and disability coverage.
  • 401(k) with a 100% employer match up to 4% of salary.
  • Paid time off, holidays, tuition reimbursement, and a $300 monthly commuter benefit.
  • Restricted stock units, cell phone reimbursement, Teladoc, Calm, and MetLife Legal benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →