Назад
1 день назад

Datacenter Infrastructure Specialist (AI)

105 000 - 140 000$
Формат работы
remote (только Europe)
Тип работы
fulltime
Грейд
senior
Страна
Europe
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Datacenter Infrastructure Specialist (AI) (GPU infrastructure/RDMA/InfiniBand): Managing the technical lifecycle and operational health of a high-density GPU fleet with an accent on hardware validation, networking performance, and distributed AI/ML workloads. Focus on automating network triage with LLMs and AI agents, troubleshooting kernel and hardware-interface issues, and coordinating incident resolution.

Datacenter Infrastructure Specialist

Company

Runpod

Conditions

6 days agoSalary: 105K - 140K

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will manage the technical lifecycle and operational health of a high-density GPU fleet. You will validate hardware deployments, monitor performance and uptime, troubleshoot networking and systems issues, automate operational workflows with AI tools, coordinate incident communications, and support infrastructure partners.

Requirements

  • 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering
  • Datacenter networking and performance troubleshooting proficiency
  • Exposure to RDMA, InfiniBand, or RoCE
  • Experience with the NVIDIA Software Stack and multi-node performance tuning
  • Linux system administration
  • Containerization with Docker
  • System-level troubleshooting and performance tuning at kernel and hardware interface layers
  • Written and verbal communication skills
  • Ability to participate in a future on-call rotation
  • Eligibility to work in the United States

Responsibilities

  • Validate new hardware and partner deployments for distributed AI and ML workloads
  • Monitor fleet health and identify performance degradation
  • Audit downtime and provide technical data to protect customer SLAs
  • Automate network triage and generate dynamic runbooks using LLMs and AI agents
  • Coordinate technical incident communications and translate outages into actionable resolutions
  • Support infrastructure partners

Benefits

  • Stock options
  • Medical, dental, and vision plans with 100% employee coverage and partial dependent coverage
  • Flexible PTO
  • $1,200 home office and equipment stipend

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -