Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Manager, Cluster Engineering & Deployment (AI): Turning delivered racks into accepted AI compute clusters through network bring-up, GPU fabric integration, validation, and production acceptance with an accent on deployment velocity, operational rigor, and defect reduction. Focus on leading concurrent cluster builds, designing RCCL performance and burn-in gates, and coordinating field teams, cabling vendors, Network Engineering, and hardware partners.
Location: Remote; candidates must be authorized to work in the United States
Company
TensorWave provides a secure, reliable cloud platform for delivering AI compute at scale.
What you will do
- Own and evolve the cluster deployment playbook, including staged bring-up, automated configuration, link and optics validation, cabling verification, and deployment fault triage.
- Lead deployment engineering across concurrent cluster builds through team leads and on-site engineers.
- Coordinate daily with Data Center Integration field teams and cabling vendors.
- Improve deployment velocity through tooling, pre-staging, and defect-source elimination.
- Manage defect feedback loops with Network Engineering, Layer One, and hardware and optics vendors.
- Define and standardize spares, test equipment, and deployment tooling requirements across sites.
Requirements
- 10+ years of experience in network deployment, cluster or HPC bring-up, or large-scale infrastructure delivery.
- Experience managing engineers in field or deployment environments and maintaining schedule accountability across multiple builds or sites.
- Hands-on fabric bring-up experience at scale, including hundreds of switches and thousands of links per deployment.
- Strong operational rigor in building and enforcing playbooks, gates, metrics, and blameless defect loops.
- Experience owning cluster validation end to end, including bandwidth and latency baselines, RCCL performance testing, burn-in criteria, and go/no-go acceptance gates.
Nice to have
- GPU cluster validation experience with NCCL or RCCL benchmarking.
- Automation experience with Python and Ansible.
- Advanced optics and link-layer debugging skills.
- Experience with acceptance testing as a commercial, revenue-linked gate.
Culture & Benefits
- Stock options.
- 100% paid medical, dental, and vision insurance for employees.
- Health Savings Account contributions and Flexible Spending Account access.
- Paid short-term and long-term disability insurance, life insurance options, and supplemental coverage.
- Flexible PTO, paid holidays, parental leave, and employee assistance programs.
- 401(k) and additional health and in-office benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →