3 часа назад
Director of AI Infrastructure (AI)
176 400 - 264 600$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Director of AI Infrastructure (AI/HPC): Overseeing high-performance computing infrastructure, including on-premises GPU clusters and orchestration across a hybrid cloud environment, with an accent on cluster performance, workload scheduling, storage architecture, and compute economics. Focus on scaling AI research infrastructure, optimizing NVIDIA GPU utilization across AWS/GCP and on-premises resources, and building reliable systems for researcher velocity.
Location: Seattle, Washington, United States; expected to work from the Seattle offices
Base salary: $176,400–$264,600 per year, plus bonus plans and long-term incentives.
Company
is a Seattle-based nonprofit AI research institute developing open-source AI research, large-scale models, data, robotics, and conservation technologies.
What you will do
- Oversee the availability, performance, and lifecycle of dense on-premises NVIDIA GPU clusters.
- Direct the strategy for Beaker and optimize workload scheduling across on-premises infrastructure and AWS/GCP cloud resources.
- Develop a long-term storage roadmap balancing high-throughput training workloads with durable, cost-effective petascale data storage.
- Manage GPU compute economics and make data-driven decisions about cloud bursting versus on-premises capacity investments.
- Partner with research teams and internal engineering groups to ensure infrastructure accelerates research rather than creating bottlenecks.
Requirements
- 12+ years of experience in infrastructure, systems engineering, or HPC, including at least 5 years leading multidisciplinary engineering teams.
- Bachelor’s degree in a related field; an advanced degree may substitute for equivalent technical experience.
- Direct experience managing large-scale NVIDIA GPU clusters and high-performance networking with InfiniBand or RoCE.
- Strong experience with Kubernetes, Slurm, or similar orchestration frameworks in hybrid-cloud environments.
- Experience with distributed filesystems such as WEKA, Ceph, or Lustre and cloud storage integration at scale.
- Proficiency in Go or Python and the ability to review architecture and code for internal tooling; U.S. work authorization is required.
Culture & Benefits
- Medical, dental, vision, employee assistance, HSA, HRA, and flexible spending account options.
- 401(k) plan, annual bonuses, and participation in a long-term incentive plan.
- Paid vacation, sick leave, personal days, family leave, and twelve paid holidays.
- Learning opportunities through Academy lectures, guest AI experts, and ongoing education support.
- Collaborative, transparent, inclusive environment with a focus on work-life balance.
- Monthly commuting or internet assistance and fitness and wellbeing reimbursement.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 часа назад
Senior Engineering Manager (AI Infrastructure)
146 880 - 220 320$
3 часа назад
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract (AI)
195 000 - 285 000$
Nscale
5 дней назад
Software Engineering Manager (AI Infrastructure)
300 000 - 350 000$
5 дней назад
Staff Engineering Manager (AI)
140 400 - 372 300$
Scale AI
2 дня назад
Sr. Director, Forward Deployed Engineering (AI)
285 200 - 356 500$
CoreWeave
19 часов назад
Manager, Technical Support Engineer (AI)
198 000 - 264 000$