Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Hardware Engineer, Server Infrastructure (Server Hardware/AI Infrastructure): Designing and operating scalable server hardware infrastructure for high-performance AI workloads with an accent on lifecycle automation, firmware management, monitoring, and platform qualification. Focus on troubleshooting hardware and firmware failures, driving root-cause resolution, supporting data center bring-up, and improving reliability across production fleets.
Location: New York, NY or Bellevue, WA, United States. The role supports data center operations, hardware technicians, and new regional infrastructure bring-up.
Company
CoreWeave provides cloud infrastructure, technology, and services for AI labs, startups, and enterprises.
What you will do
- Design, develop, and optimize server hardware infrastructure for high-performance workloads.
- Automate the complete server hardware lifecycle, including provisioning, configuration, firmware management, monitoring, and decommissioning.
- Build hardware and firmware management services, monitoring, alerting, runbooks, and self-service tooling for reliable operations at scale.
- Lead hardware escalations, deep troubleshooting, root-cause analysis, incident response, and long-term corrective actions.
- Evaluate, qualify, and deploy new server platforms with vendors and OEMs, including support for firmware, quality, RMA, and data center region bring-up.
- Document designs and procedures, support operations teams and technicians, and use production failures to improve automation, telemetry, and platform design.
Requirements
- Deep understanding of server hardware, components, and management technologies.
- Hands-on proficiency with Ansible or Python and programmatic interaction with server BMCs through Redfish or IPMI; Redfish is preferred.
- Experience evaluating, qualifying, deploying, and troubleshooting production server infrastructure with hardware vendors and OEMs.
- Experience participating in on-call or escalation rotations and driving incidents through root-cause resolution.
- Strong technical documentation, analytical, problem-solving, and cross-functional communication skills.
- Excellent written and verbal English communication skills are required.
Nice to have
- Experience with new data center region bring-up, GPU platforms, or rack-scale systems such as NVIDIA GB200 or GB300.
- Fleet-scale Redfish or IPMI, firmware lifecycle management, or hardware qualification experience.
- Strong Linux administration and debugging skills at fleet scale.
- Experience with Kubernetes, Prometheus, Grafana, and distributed production environments.
- Experience creating alerting, runbooks, and self-service tooling for operations teams or supporting external technical partners.
Culture & Benefits
- Full-time US benefits include fully paid medical, dental, and vision insurance.
- Life, disability, family-forming, parental leave, mental wellness, childcare, and flexible spending benefits are available.
- Benefits include tuition reimbursement, employee stock purchase program participation, 401(k) with employer match, and flexible PTO.
- Office and data center locations provide catered lunches and a casual work environment.
- The role includes an on-call rotation and collaboration across engineering, operations, vendors, and customer-facing stakeholders.
- Access to export-controlled information requires US-person status, eligibility without export authorization, or eligibility and reasonable likelihood of obtaining the required authorization.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Roblox
5 дней назад
Senior Hardware Engineer - Infrastructure
243 290 - 295 250$
1 день назад
Lead Field Service Engineer (Hardware)
5 дней назад
CAD Engineer – Infrastructure (ASIC/EDA)
230 773 - 323 082$
5 дней назад
Hardware Engineer – Host Processor Systems (Embedded Hardware)
103 500 - 191 900$
5 дней назад
Principal Hardware Engineer (Networking)
153 500 - 310 500$
Nebius
1 день назад
Field Data Center Hardware Engineer (AI)
112 000 - 140 000$