Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Field Deploy Engineer (AI Data Centers): Leading on-site deployment of servers, storage, networking, and supporting systems for new data center locations with an accent on hardware installation, infrastructure readiness, and multi-stakeholder coordination. Focus on diagnosing hardware, firmware, and Linux issues, validating operational readiness, and building repeatable deployment practices for GPU infrastructure.
Location: Singapore; on-site data center work and frequent travel required
Company
Nebius builds a full-stack AI cloud platform covering data and model training through production deployment, with infrastructure spanning compute, storage, networking, and applied AI.
What you will do
- Own hardware deployment plans for new data center sites, aligning scope, timelines, and resources with project stakeholders.
- Coordinate logistics, supply chain, construction, networking, vendors, contractors, and local staff for GPU server deployments.
- Supervise racking, cabling, powering, configuration, firmware upgrades, and integration with management systems.
- Perform functional checks, acceptance testing, troubleshooting, and readiness validation for deployed infrastructure.
- Document installation procedures, asset data, lessons learned, and deployment best practices.
- Support the transition of deployed hardware to operations teams and provide status updates to project leadership.
Requirements
- 3+ years of experience in data center, hardware, or IT field delivery roles.
- Strong knowledge of server, storage, and network installation, including racking, cabling, power, and cooling.
- Hands-on experience working on data center floors with server hardware and troubleshooting Linux.
- Experience coordinating deployments across multiple stakeholders and vendors.
- Ability to interpret technical documentation, layouts, and wiring diagrams.
- Ability to diagnose hardware, BIOS/BMC firmware, and Linux issues, including logs, drivers, storage, and performance.
Nice to have
- Experience with GPU server platforms and tools such as nvidia-smi and dcgmi.
- Knowledge of Redfish, BMC tooling, ipmitool, firmware lifecycle management, and staged rollouts.
- Bash and basic Python scripting for log collection, triage automation, and reliability analysis.
- Exposure to OCP platforms, ODM manufacturing ecosystems, asset tracking, change management, or quality assurance.
- Understanding of health, safety, and environmental standards in data center or industrial settings.
Culture & Benefits
- Competitive compensation and opportunities for career growth and learning.
- Flexibility, ownership, and a collaborative working environment.
- Opportunity to contribute to impactful AI infrastructure projects.
- International environment with experienced engineering and operations specialists.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →