Назад
Company hidden
5 часов назад

Operations Engineer

208 000 - 269 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Operations Engineer (Engineering): Building and operating large-scale GPU compute infrastructure with an accent on automation, hardware lifecycle management, and fleet health. Focus on designing repair pipelines, GPU qualification platforms, and orchestration tooling to ensure reliability and scalability of hyperscale AI compute.

Location

Location: San Francisco, CA; Austin, TX; New York, NY; Seattle, WA (On-site)

Salary

Salary: $208,000 – $269,000 per year plus equity

Company

hirify.global builds civilization-scale infrastructure for AI, focusing on delivering massive GPU compute capacity with speed and scale as key differentiators.

What you will do

  • Build and automate repair pipelines for GPU fleet management to handle failures without human intervention.
  • Qualify new GPU generations rapidly to define production readiness before site launches.
  • Migrate live compute clusters and manage Kubernetes-orchestrated bare metal infrastructure at scale.
  • Develop observability and orchestration layers for real-time fleet health and performance monitoring.
  • Design and maintain metrics pipelines, alerting systems, and unified health views for compute fleet.
  • Own firmware-level telemetry tooling such as Redfish and BMC for repair automation and health monitoring.

Requirements

  • Location: Must work on-site in one of the specified US cities
  • Strong hardware instincts including firmware and silicon-level failure reasoning.
  • Ability to handle ambiguity and learn quickly in unfamiliar domains.
  • Experience with AI tooling such as LLM APIs and agentic frameworks.
  • Proven track record shipping production automation in any programming language.
  • Incident management experience including running pagers and postmortems.

Nice to have

  • Experience with hardware lifecycle management and RMA automation.
  • Familiarity with BMC/Redfish or IPMI tooling.
  • Knowledge of GPU qualification or burn-in frameworks.
  • Experience with workflow and orchestration engines like Temporal or Cadence.
  • Skills in metrics and alerting pipelines using Prometheus and Grafana.
  • Programming experience in Go or Python.

Culture & Benefits

  • Competitive total compensation including salary and equity.
  • Retirement or pension plans aligned with local norms.
  • Health, dental, and vision insurance.
  • Generous paid time off policy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →