Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Technical Writer (Infrastructure): Building and maintaining the L3 knowledge system for GPU platforms and Linux diagnostics with an accent on translating complex engineering investigations into operational procedures. Focus on establishing documentation governance, validating runbooks with end-users, and ensuring production readiness for new hardware.
Location: Amsterdam, Netherlands (Willingness to travel regularly to data center locations in EMEA is required)
Company
Nebius is building a full-stack AI cloud platform designed to support developers and enterprises with large-scale GPU orchestration and inference optimization.
What you will do
- Build and maintain the L3 knowledge base for GPU server platforms, server hardware, firmware, and Linux-level diagnostics.
- Translate deep technical investigations from R&D engineers into clear, repeatable runbooks, SOPs, and troubleshooting guides.
- Define and enforce documentation governance, including templates, quality standards, ownership rules, and review cycles.
- Validate procedures end-to-end by testing them with L1 and L2 technicians and improving content based on feedback.
- Create complete documentation packages to ensure new hardware platforms are ready for production support.
- Travel to EMEA data centers to observe real-world procedures and capture operational knowledge.
Requirements
- Hands-on experience in data center, server infrastructure, production operations, or SRE.
- Working knowledge of Linux, server hardware, firmware, and OOB management (IPMI, BMC, OpenBMC, or Redfish).
- Proven track record of creating operational runbooks or SOPs used successfully by operations teams.
- Ability to write clear, precise, and structured English for readers with varying technical experience.
- Must be authorized to work in the country of application (Netherlands).
- Willingness to travel regularly to data center locations in EMEA.
Nice to have
- Experience with NVIDIA GPU platforms and tools (nvidia-smi, DCGM, dcgmi).
- Exposure to HGX or OCP-based platforms and ODM manufacturing ecosystems.
- Proficiency in Bash or Python for log collection and diagnostics.
- Experience with documentation-as-code, Git-based workflows, or large-scale knowledge bases.
Culture & Benefits
- Competitive compensation and opportunities for career growth and learning.
- Collaborative and innovative culture within an international team.
- High degree of flexibility and ownership over the documentation product.
- Opportunity to work on high-impact AI infrastructure projects.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →