Назад
12 часов назад

Data Center Engineer, Reliability & Infrastructure Management – Compute Supply

320 000 - 405 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Center Engineer, Reliability & Infrastructure Management – Compute Supply (AI infrastructure): Building power, availability, topology, and load-management models for large-scale AI compute infrastructure with an accent on data center electrical and mechanical systems, capacity planning, and failure management. Focus on validating telemetry, defining failure domains, forecasting workload power demand, and driving SLA-grade improvements with data center and cloud partners.

Location: San Francisco, CA or New York City, NY; hybrid policy requiring staff to be in one of the offices at least 25% of the time

Annual salary: $320,000–$405,000 USD

Company

Anthropic develops reliable, interpretable, and steerable AI systems designed to be safe and beneficial.

What you will do

  • Own power and cooling topology data for the compute fleet, validate telemetry against building behavior, and define load-management requirements through commissioning and incident support.
  • Build and audit availability models for data center electrical and mechanical systems and define cloud availability zones and failure domains.
  • Model power draw from chip to rack to facility by workload and hardware generation to support capacity planning, design envelopes, and load forecasting.
  • Work with data center developers, operators, cloud providers, and chip vendors on architecture reviews, reliability studies, rack power specifications, and operator interfaces.
  • Translate infrastructure requirements across hardware, software, facilities, legal, commercial, and external partner teams.

Requirements

  • Bachelor’s degree or equivalent experience in electrical engineering, mechanical engineering, power systems, reliability engineering, controls engineering, or a related field.
  • 5+ years of experience in data center infrastructure, facility engineering, or reliability engineering.
  • Deep knowledge of data center power distribution, cooling architectures, redundancy schemes, and infrastructure failure-mode management.
  • Experience building quantitative reliability or availability models, power or capacity models and forecasts, or software-based power management, load-shedding, or control systems.
  • Familiarity with SCADA, BMS, EPMS, telemetry pipelines, control systems, and software bridging IT and OT.
  • Track record of cross-functional collaboration across hardware, software, and facilities teams.

Nice to have

  • Experience with accelerator-class deployments, rack power architectures, and power-management interfaces.
  • Experience with integrated systems testing and commissioning at L4/L5 or writing sequences of operation.
  • Experience with SLA development, availability commitments, or service-credit frameworks.
  • Experience with energy storage, microgrids, demand response, or behind-the-meter generation.
  • Exposure to ML or optimization techniques applied to infrastructure or energy systems.

Culture & Benefits

  • Collaborative environment spanning research, engineering, policy, and business disciplines.
  • Flexible working hours and office-based collaboration.
  • Generous vacation and parental leave.
  • Competitive compensation, benefits, and optional equity donation matching.
  • Visa sponsorship is available, with individual role and candidate eligibility assessed during the offer process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →