Назад
14 часов назад

Repairs Program Lead - Data Center Operations

320 000 - 405 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Repairs Program Lead - Data Center Operations (Data Center Hardware): Defining and managing end-to-end hardware repair programs across a growing fleet of data centers with an accent on repair SLAs, break-fix operations, RMA and reverse logistics, and spares planning. Focus on analyzing failure patterns, driving corrective actions with engineering and suppliers, and maintaining compute availability across multiple sites.

Location: Remote-friendly with travel required; San Francisco, California. Staff are expected to work from one of the offices at least 25% of the time.

Annual salary: $320,000–$405,000 USD

Company

Anthropic develops reliable, interpretable, and steerable AI systems intended to be safe and beneficial for users and society.

What you will do

  • Define the global repair strategy, including SLAs, prioritization rules, escalation paths, reporting, and standards across all data center sites.
  • Own repair turnaround time, backlog, fleet availability impact, and repair dashboards based on ticket and telemetry data.
  • Develop procedures for triage, break-fix work, return-to-service validation, and train site operations partners.
  • Manage RMA, warranty, reverse logistics, and depot repair programs with OEMs, ODMs, and repair vendors.
  • Set spares pool sizing and inventory levels in coordination with supply chain and asset management.
  • Analyze hardware failures, lead operating reviews, and drive corrective actions with engineering, suppliers, vendors, and site operations.

Requirements

  • 8+ years of data center operations experience as a manager, technical lead, or in a related role, with accountability for production availability.
  • Proven experience running large-scale break-fix programs across multiple sites.
  • Experience managing vendors, OEMs, ODMs, or contract workforces against measurable SLAs and corrective actions.
  • Hands-on technical depth in server, network, and rack-level hardware to verify repair quality and audit vendor claims.
  • Experience building or substantially improving operational processes and using ticket, telemetry, and inventory data for decisions.
  • A bachelor's degree in a relevant field or equivalent practical experience.

Nice to have

  • Experience with GPU or accelerator systems and high-density liquid-cooled infrastructure, including tray, cold plate, and manifold-level repairs.
  • Experience with hyperscale OEM/ODM RMA and warranty programs, failure analysis, and supplier quality engagement.
  • Experience with spares planning, reverse logistics, or depot repair at data center scale.
  • Experience delivering repairs in partner-operated or colocation sites staffed by third parties.
  • Familiarity with optics and high-speed interconnect failures.

Culture & Benefits

  • Collaborative work across research, engineering, policy, and business functions.
  • Flexible working hours and an office environment designed for collaboration.
  • Competitive compensation, optional equity donation matching, generous vacation, and parental leave.
  • Visa sponsorship is available, subject to role and candidate eligibility.
  • Anthropic is a public benefit corporation headquartered in San Francisco.

Hiring process

  • Applicants should review Anthropic's guidance on using AI during the application process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →