Назад
3 дня назад

Operations Manager (AI Infrastructure)

225 000 - 235 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Operations Manager (AI Infrastructure): Managing the operational and analytical supply side of a GPU fleet across neocloud and bare metal environments with an accent on fleet health, capacity reconciliation, supplier accountability, and service-level performance. Focus on monitoring contracted versus provisioned versus healthy capacity, driving replacement and repair SLAs, enforcing supplier remedies, and coordinating maintenance across infrastructure, finance, legal, and security.

Location: San Francisco, United States; hybrid

Salary: $225,000–$235,000 per year plus equity

Company

Baseten provides mission-critical AI inference infrastructure, applied AI research, and developer tooling for companies bringing machine learning models into production.

What you will do

  • Manage the operational and analytical supply side of the GPU fleet across neocloud and bare metal environments.
  • Drive suppliers to keep the maximum amount of GPU capacity online and healthy.
  • Reconcile contracted, provisioned, healthy, and utilized capacity by supplier and cluster.
  • Own supplier accountability for replacement SLAs, mean time to repair, RMA cycles, and fleet health.
  • Monitor SLA performance, pursue credit claims, and enforce remediation plans.
  • Coordinate internal communications for supplier maintenance affecting availability.

Requirements

  • 5–10+ years of infrastructure experience focused on the compute lifecycle and maximizing functional compute.
  • Direct experience managing GPU, server, or data center hardware supplier relationships.
  • Understanding of fleet health, RMA processes, and the differences between contracted and delivered capacity.
  • Strong analytical skills, including independently extracting data, building reports, and generating insights.
  • Ability to operate effectively in ambiguous environments and define needed processes.
  • Strong collaboration skills across finance, infrastructure and engineering, legal, and security.

Nice to have

  • Experience at a hyperscaler, neocloud provider, or AI infrastructure company.
  • Familiarity with NVIDIA H100, H200, or GB200-class systems, power and thermal constraints, and compute supply chains.
  • Experience running formal supplier corrective actions.

Culture & Benefits

  • Competitive compensation with meaningful equity.
  • Medical, dental, and vision insurance fully covered for employees and dependents.
  • Flexible PTO and a company-wide winter break.
  • Paid parental leave and a fertility and family-building stipend.
  • Company-facilitated 401(k).
  • Exposure to a variety of machine learning startups and AI infrastructure environments.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →