3 часа назад
Senior Engineering Manager (AI Infrastructure)
146 880 - 220 320$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Engineering Manager (AI Infrastructure): Operating and improving high-performance computing infrastructure for AI research with an accent on on-premise GPU clusters, hybrid-cloud orchestration, storage, and resource allocation. Focus on leading systems engineers and SREs, optimizing Beaker scheduling, managing petascale storage, and maintaining reliable infrastructure for large-scale model training.
Location: Seattle, Washington, United States; expected to work from the Seattle offices
Base salary: $146,880–$220,320 per year, plus bonus plans and a long-term incentive plan.
Company
is a Seattle-based nonprofit research institute focused on open-source AI development, foundational research, large-scale models, data, robotics, and conservation.
What you will do
- Lead day-to-day operations, availability, performance, and health of dense on-premise NVIDIA GPU clusters.
- Operate and improve Beaker orchestration across on-premise infrastructure and AWS/GCP cloud resources.
- Manage distributed storage for high-throughput model training and durable petascale research data.
- Allocate GPU resources against budget, track utilization, and recommend cloud bursting or on-premise capacity investments.
- Partner with research teams to keep infrastructure reliable and responsive to their objectives.
- Manage and grow a team of systems engineers, SREs, and software developers.
Requirements
- 12+ years of experience in infrastructure, systems engineering, or HPC, or an advanced degree with 8+ years of technical experience.
- At least 2 years supervising a team of 5 or more engineers.
- Bachelor's degree in a related field, with a relevant advanced degree accepted as a substitute for equivalent technical experience.
- Direct experience operating large-scale NVIDIA GPU clusters and InfiniBand or RoCE networks.
- Strong experience with Kubernetes, Slurm, or similar orchestration frameworks in hybrid-cloud environments.
- Hands-on distributed filesystem and cloud storage experience, plus proficiency in Go or Python and SDLC processes.
Culture & Benefits
- Open-science environment focused on releasing research, data, code, and models to the scientific community.
- Medical, dental, vision, employee assistance, HSA, HRA, and flexible spending account plans.
- 401(k) plan, annual bonuses, and long-term incentive plan.
- Paid vacation, sick leave, personal days, holidays, and family leave.
- Monthly commuting or internet and fitness and wellbeing allowances.
- Learning opportunities through Ai2 Academy lectures, guest experts, and continuing education support.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 часа назад
Director of AI Infrastructure (AI)
176 400 - 264 600$
3 часа назад
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract (AI)
195 000 - 285 000$
2 дня назад
Senior Manager, Agentic AI Integrations
174 051 - 273 775$
5 дней назад
Staff Engineering Manager (AI)
140 400 - 372 300$
2 часа назад
Senior Manager, Infrastructure Platform Engineering (AI)
245 000 - 295 000$
Nscale
5 дней назад
Software Engineering Manager (AI Infrastructure)
300 000 - 350 000$