IT Infrastructure Engineer – RMA & Hardware Diagnostics (AI Cloud)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
IT Infrastructure Engineer – RMA & Hardware Diagnostics (AI Cloud): Owning advanced hardware troubleshooting and RMA lifecycle management within production data center environments with an accent on root cause analysis and vendor coordination. Focus on performing deep diagnostics across enterprise server platforms, validating failed components, and improving fleet-wide reliability.
Location: On-site in Kansas City, Missouri, United States
Salary: $112,700 - $140,800 OTE
Company
Nebius is building a full-stack AI cloud platform to support developers and enterprises from data and model training through to production deployment.
What you will do
- Perform advanced firmware and hardware diagnostics on enterprise server platforms, including CPU, memory, PCIe devices, GPUs, and storage subsystems.
- Troubleshoot complex hardware failures using system logs, BMC/IPMI interfaces, and BIOS diagnostics.
- Own the full RMA lifecycle, including validation of failed components, warranty claim creation, and resolution with OEM vendors.
- Conduct structured root cause analysis and document findings to prevent repeat failures.
- Develop and standardize diagnostic playbooks, troubleshooting workflows, and hardware validation procedures.
- Collaborate with data center operations and engineering teams to reduce MTTR and improve overall fleet reliability.
Requirements
- 5+ years of hands-on experience with enterprise server hardware in a production data center environment.
- Deep understanding of x86 server architecture, including CPUs, memory, PCIe devices, GPUs, and power subsystems.
- Strong experience performing firmware and BIOS/BMC diagnostics and upgrades.
- Advanced Linux command-line troubleshooting skills, including log analysis and hardware-level diagnostics.
- Proven experience managing hardware RMA processes and working directly with OEM vendors.
- High proficiency in spoken and written English.
Nice to have
- Experience performing board-level diagnostics and component-level repair (SMD rework).
- Familiarity with data center networking equipment and basic network troubleshooting.
- Experience supporting GPU-dense or high-performance compute environments.
- Valid driver’s license.
Culture & Benefits
- Competitive compensation and career growth opportunities.
- Collaborative and innovative culture with the opportunity to work on impactful AI projects.
- International environment with highly talented engineering teams.
- Emphasis on flexibility, trust, and real ownership within a fast-moving growth environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →