Назад
Company hidden
3 часа назад

Senior Engineering Manager (AI Infrastructure)

146 880 - 220 320$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Engineering Manager (AI Infrastructure): Operating and improving high-performance computing infrastructure for AI research with an accent on on-premise GPU clusters, hybrid-cloud orchestration, storage, and resource allocation. Focus on leading systems engineers and SREs, optimizing Beaker scheduling, managing petascale storage, and maintaining reliable infrastructure for large-scale model training.

Location: Seattle, Washington, United States; expected to work from the Seattle offices

Base salary: $146,880–$220,320 per year, plus bonus plans and a long-term incentive plan.

Company

hirify.global is a Seattle-based nonprofit research institute focused on open-source AI development, foundational research, large-scale models, data, robotics, and conservation.

What you will do

  • Lead day-to-day operations, availability, performance, and health of dense on-premise NVIDIA GPU clusters.
  • Operate and improve Beaker orchestration across on-premise infrastructure and AWS/GCP cloud resources.
  • Manage distributed storage for high-throughput model training and durable petascale research data.
  • Allocate GPU resources against budget, track utilization, and recommend cloud bursting or on-premise capacity investments.
  • Partner with research teams to keep infrastructure reliable and responsive to their objectives.
  • Manage and grow a team of systems engineers, SREs, and software developers.

Requirements

  • 12+ years of experience in infrastructure, systems engineering, or HPC, or an advanced degree with 8+ years of technical experience.
  • At least 2 years supervising a team of 5 or more engineers.
  • Bachelor's degree in a related field, with a relevant advanced degree accepted as a substitute for equivalent technical experience.
  • Direct experience operating large-scale NVIDIA GPU clusters and InfiniBand or RoCE networks.
  • Strong experience with Kubernetes, Slurm, or similar orchestration frameworks in hybrid-cloud environments.
  • Hands-on distributed filesystem and cloud storage experience, plus proficiency in Go or Python and SDLC processes.

Culture & Benefits

  • Open-science environment focused on releasing research, data, code, and models to the scientific community.
  • Medical, dental, vision, employee assistance, HSA, HRA, and flexible spending account plans.
  • 401(k) plan, annual bonuses, and long-term incentive plan.
  • Paid vacation, sick leave, personal days, holidays, and family leave.
  • Monthly commuting or internet and fitness and wellbeing allowances.
  • Learning opportunities through Ai2 Academy lectures, guest experts, and continuing education support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →