Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
L3 Support Engineer (Linux) (AI cloud infrastructure): Building a datacenter L3 support capability for servers, BIOS/BMC firmware, and deep Linux diagnostics across APAC with an accent on root cause analysis, GPU failures, and hardware-software interactions. Focus on detecting cross-site incident patterns, driving permanent fixes with R&D and ODM vendors, and creating scalable runbooks for L1/L2 enablement.
Location: Singapore. The role supports datacenters across APAC and may require travel to datacenters for complex troubleshooting, platform readiness, or incident containment.
Company
Nebius is building a full-stack AI cloud platform for data processing, model training, and production deployment, with infrastructure spanning compute, storage, networking, and applied AI.
What you will do
- Lead deep root cause analysis for GPU failures, firmware issues, Linux faults, and hardware-software interactions.
- Identify recurring issues across datacenter sites and convert investigations into durable platform fixes.
- Own technical workstreams during high-severity incidents and drive evidence-based escalations with ODM vendors and R&D.
- Support BIOS/BMC firmware validation, risk assessment, staged rollouts, and rollback planning.
- Create runbooks, troubleshooting guides, error catalogs, and playbooks that improve L1/L2 support capabilities.
- Travel to datacenters when required for complex troubleshooting, new platform readiness, or incident containment.
Requirements
- Hands-on experience with datacenter servers and deep Linux troubleshooting.
- Ability to diagnose hardware, BIOS/BMC firmware, Linux logs and drivers, storage fundamentals, and performance issues.
- Experience with structured incident response and communication under pressure.
- Experience driving evidence-based escalations with vendors or R&D teams.
- Fluent English, written and spoken.
- Applicants must be authorized to work in Singapore and provide proof of employment eligibility.
Nice to have
- Experience with GPU server platforms and tools such as nvidia-smi, dcgmi, and Linux log correlation.
- Familiarity with ipmitool, Redfish, firmware lifecycle management, and staged rollouts.
- Bash and basic Python scripting for log collection, triage automation, and reliability analysis.
- Exposure to OCP platforms, ODM manufacturing ecosystems, or enterprise bare-metal customers under contractual SLAs.
Culture & Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility, ownership, and a collaborative working environment.
- Opportunity to contribute to impactful AI infrastructure projects.
- International environment with a focus on meaningful impact and continuous growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Linux Systems / Infrastructure Engineer (RHEL)
Adyen
8 дней назад
Senior Linux Infrastructure Engineer
7 дней назад
Associate Linux Engineer, Technology II (Biotech)
65 500 - 125 500$
4 дня назад
NOC Engineer (Network Operations)
9 дней назад
Data Center Site Manager / Supervisor (AI/HPC)
9 дней назад