Назад
15 дней назад

Senior Manager, Cluster Engineering & Deployment (AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Manager, Cluster Engineering & Deployment (AI): Turning delivered racks into accepted AI compute clusters through network bring-up, GPU fabric integration, validation, and production acceptance with an accent on deployment velocity, operational rigor, and defect reduction. Focus on leading concurrent cluster builds, designing RCCL performance and burn-in gates, and coordinating field teams, cabling vendors, Network Engineering, and hardware partners.

Location: Remote; candidates must be authorized to work in the United States

Company

TensorWave provides a secure, reliable cloud platform for delivering AI compute at scale.

What you will do

  • Own and evolve the cluster deployment playbook, including staged bring-up, automated configuration, link and optics validation, cabling verification, and deployment fault triage.
  • Lead deployment engineering across concurrent cluster builds through team leads and on-site engineers.
  • Coordinate daily with Data Center Integration field teams and cabling vendors.
  • Improve deployment velocity through tooling, pre-staging, and defect-source elimination.
  • Manage defect feedback loops with Network Engineering, Layer One, and hardware and optics vendors.
  • Define and standardize spares, test equipment, and deployment tooling requirements across sites.

Requirements

  • 10+ years of experience in network deployment, cluster or HPC bring-up, or large-scale infrastructure delivery.
  • Experience managing engineers in field or deployment environments and maintaining schedule accountability across multiple builds or sites.
  • Hands-on fabric bring-up experience at scale, including hundreds of switches and thousands of links per deployment.
  • Strong operational rigor in building and enforcing playbooks, gates, metrics, and blameless defect loops.
  • Experience owning cluster validation end to end, including bandwidth and latency baselines, RCCL performance testing, burn-in criteria, and go/no-go acceptance gates.

Nice to have

  • GPU cluster validation experience with NCCL or RCCL benchmarking.
  • Automation experience with Python and Ansible.
  • Advanced optics and link-layer debugging skills.
  • Experience with acceptance testing as a commercial, revenue-linked gate.

Culture & Benefits

  • Stock options.
  • 100% paid medical, dental, and vision insurance for employees.
  • Health Savings Account contributions and Flexible Spending Account access.
  • Paid short-term and long-term disability insurance, life insurance options, and supplemental coverage.
  • Flexible PTO, paid holidays, parental leave, and employee assistance programs.
  • 401(k) and additional health and in-office benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →