Назад
1 день назад

Senior Hardware Engineer, Server Infrastructure

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Hardware Engineer, Server Infrastructure (Server Hardware/AI Infrastructure): Designing and operating scalable server hardware infrastructure for high-performance AI workloads with an accent on lifecycle automation, firmware management, monitoring, and platform qualification. Focus on troubleshooting hardware and firmware failures, driving root-cause resolution, supporting data center bring-up, and improving reliability across production fleets.

Location: New York, NY or Bellevue, WA, United States. The role supports data center operations, hardware technicians, and new regional infrastructure bring-up.

Company

CoreWeave provides cloud infrastructure, technology, and services for AI labs, startups, and enterprises.

What you will do

  • Design, develop, and optimize server hardware infrastructure for high-performance workloads.
  • Automate the complete server hardware lifecycle, including provisioning, configuration, firmware management, monitoring, and decommissioning.
  • Build hardware and firmware management services, monitoring, alerting, runbooks, and self-service tooling for reliable operations at scale.
  • Lead hardware escalations, deep troubleshooting, root-cause analysis, incident response, and long-term corrective actions.
  • Evaluate, qualify, and deploy new server platforms with vendors and OEMs, including support for firmware, quality, RMA, and data center region bring-up.
  • Document designs and procedures, support operations teams and technicians, and use production failures to improve automation, telemetry, and platform design.

Requirements

  • Deep understanding of server hardware, components, and management technologies.
  • Hands-on proficiency with Ansible or Python and programmatic interaction with server BMCs through Redfish or IPMI; Redfish is preferred.
  • Experience evaluating, qualifying, deploying, and troubleshooting production server infrastructure with hardware vendors and OEMs.
  • Experience participating in on-call or escalation rotations and driving incidents through root-cause resolution.
  • Strong technical documentation, analytical, problem-solving, and cross-functional communication skills.
  • Excellent written and verbal English communication skills are required.

Nice to have

  • Experience with new data center region bring-up, GPU platforms, or rack-scale systems such as NVIDIA GB200 or GB300.
  • Fleet-scale Redfish or IPMI, firmware lifecycle management, or hardware qualification experience.
  • Strong Linux administration and debugging skills at fleet scale.
  • Experience with Kubernetes, Prometheus, Grafana, and distributed production environments.
  • Experience creating alerting, runbooks, and self-service tooling for operations teams or supporting external technical partners.

Culture & Benefits

  • Full-time US benefits include fully paid medical, dental, and vision insurance.
  • Life, disability, family-forming, parental leave, mental wellness, childcare, and flexible spending benefits are available.
  • Benefits include tuition reimbursement, employee stock purchase program participation, 401(k) with employer match, and flexible PTO.
  • Office and data center locations provide catered lunches and a casual work environment.
  • The role includes an on-call rotation and collaboration across engineering, operations, vendors, and customer-facing stakeholders.
  • Access to export-controlled information requires US-person status, eligibility without export authorization, or eligibility and reasonable likelihood of obtaining the required authorization.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →