Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
L3 Support Engineer (SOPs and Runbooks) (Datacenter Infrastructure/AI Cloud): Own the L3 knowledge system for GPU platforms, server hardware, firmware, out-of-band management, and Linux diagnostics with an accent on SOP governance, runbook quality, and incident-driven documentation. Focus on converting L3 investigations into tested procedures, validating them with L1 and L2 technicians, and preparing documentation for new hardware platforms.
Location: Amsterdam, Netherlands; willingness to travel to datacenters is required.
Company
Nebius builds a full-stack AI cloud platform for data processing, model training, inference, and production deployment.
What you will do
- Own the L3 documentation standard, including templates, quality criteria, ownership, approval workflows, and review cycles.
- Build and maintain SOPs, runbooks, troubleshooting guides, and error-code libraries for GPU platforms, server hardware, firmware, out-of-band management, and Linux diagnostics.
- Join L3 and R&D investigations, capture symptoms, root causes, evidence, resolutions, and preventive actions, then convert findings into repeatable procedures.
- Validate procedures end to end with L1 and L2 technicians and improve documentation when escalations recur.
- Prepare documentation packages for new hardware platforms before production support and visit datacenters to observe procedures and capture incident knowledge.
Requirements
- Hands-on experience with datacenter servers and Linux troubleshooting.
- Working knowledge of server hardware, BIOS/BMC firmware, and out-of-band management such as IPMI or Redfish.
- Demonstrated experience writing runbooks, SOPs, or troubleshooting guides used by operations teams.
- Strong working knowledge of Atlassian Confluence.
- Clear, structured written English for readers with different technical levels; fluent written and spoken English is required.
- Applicants must be authorized to work in the country in which they apply; willingness to travel is required.
Nice to have
- Experience with GPU server tooling such as nvidia-smi, dcgmi, and Linux log correlation.
- Experience with ipmitool, Redfish workflows, and firmware lifecycle management.
- Bash and basic Python scripting for log collection.
- Documentation-as-code or Git-based workflow experience.
- Exposure to OCP-based platforms and ODM ecosystems.
Culture & Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative, innovative, and international environment.
- Opportunity to work on impactful AI projects.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →