Назад
3 дня назад

L3 Support Engineer (Linux)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
UK/Singapore/US +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
L3 Support Engineer (Linux) (AI cloud infrastructure): Building a datacenter L3 support capability for servers, BIOS/BMC firmware, and deep Linux diagnostics across APAC with an accent on root cause analysis, GPU failures, and hardware-software interactions. Focus on detecting cross-site incident patterns, driving permanent fixes with R&D and ODM vendors, and creating scalable runbooks for L1/L2 enablement.

Location: Singapore. The role supports datacenters across APAC and may require travel to datacenters for complex troubleshooting, platform readiness, or incident containment.

Company

Nebius is building a full-stack AI cloud platform for data processing, model training, and production deployment, with infrastructure spanning compute, storage, networking, and applied AI.

What you will do

  • Lead deep root cause analysis for GPU failures, firmware issues, Linux faults, and hardware-software interactions.
  • Identify recurring issues across datacenter sites and convert investigations into durable platform fixes.
  • Own technical workstreams during high-severity incidents and drive evidence-based escalations with ODM vendors and R&D.
  • Support BIOS/BMC firmware validation, risk assessment, staged rollouts, and rollback planning.
  • Create runbooks, troubleshooting guides, error catalogs, and playbooks that improve L1/L2 support capabilities.
  • Travel to datacenters when required for complex troubleshooting, new platform readiness, or incident containment.

Requirements

  • Hands-on experience with datacenter servers and deep Linux troubleshooting.
  • Ability to diagnose hardware, BIOS/BMC firmware, Linux logs and drivers, storage fundamentals, and performance issues.
  • Experience with structured incident response and communication under pressure.
  • Experience driving evidence-based escalations with vendors or R&D teams.
  • Fluent English, written and spoken.
  • Applicants must be authorized to work in Singapore and provide proof of employment eligibility.

Nice to have

  • Experience with GPU server platforms and tools such as nvidia-smi, dcgmi, and Linux log correlation.
  • Familiarity with ipmitool, Redfish, firmware lifecycle management, and staged rollouts.
  • Bash and basic Python scripting for log collection, triage automation, and reliability analysis.
  • Exposure to OCP platforms, ODM manufacturing ecosystems, or enterprise bare-metal customers under contractual SLAs.

Culture & Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility, ownership, and a collaborative working environment.
  • Opportunity to contribute to impactful AI infrastructure projects.
  • International environment with a focus on meaningful impact and continuous growth.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →