Назад
2 дня назад

Engineering Manager (AI Infrastructure)

330 000 - 440 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Engineering Manager (AI Infrastructure): Leading teams responsible for deploying and operating production GPU fleet infrastructure with an accent on reliability, bare-metal provisioning, orchestration, and datacenter systems. Focus on managing large-scale compute environments, driving cross-functional deployments, improving automation and efficiency, and maintaining reliable systems with real SLAs.

Location: Hybrid, with presence in the San Francisco, San Jose, or Bellevue office 4 days per week; Tuesday is the designated work-from-home day.

Annual salary: $330,000–$440,000 for San Francisco or San Jose; $297,000–$396,000 for Bellevue.

Company

AI cloud infrastructure company building large-scale GPU compute systems for researchers, enterprises, and hyperscalers.

What you will do

  • Lead and grow distributed engineering teams responsible for production systems infrastructure.
  • Coordinate cross-functional projects and deployments, aligning stakeholders and delivering against deadlines.
  • Improve efficiency through tooling, process optimization, and automation.
  • Provide visibility into project progress, risks, outcomes, staffing, priorities, and deliverables.
  • Participate in new technology qualification, incident management, and incident review programs.
  • Conduct 1:1s, provide feedback, and support team members’ career development.

Requirements

  • 3+ years of experience leading or managing engineers in AI/ML infrastructure or another large-scale compute environment.
  • Experience owning production systems with real SLAs and balancing operational reliability with long-term technical improvements.
  • Confidence working in Linux and debugging across operating system, hardware, and networking layers.
  • Ability to lead technical design for medium-to-large initiatives, drive alignment, and deliver solutions.
  • Experience building high-performing teams through hiring, upskilling, skills planning, performance management, and clear expectations.
  • Strong problem-solving and troubleshooting skills, with interest in the intersection of hardware, software, and physical datacenter infrastructure.

Nice to have

  • Linux systems administration, TCP/IP networking, automation, and scripting.
  • Bare-metal provisioning and lifecycle management with PXE, Redfish, IPMI, BMC, DHCP, or DNS.
  • Strong coding ability, APIs, distributed systems, and automation pipelines.
  • GPU acceleration, virtualization, cloud computing, InfiniBand, racks, switches, and power domains.
  • NetBox or similar source-of-truth and DCIM tooling, Linux distribution building, OS customization, or imaging.

Culture & Benefits

  • Generous cash and equity compensation.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan with a 2% company match for US employees.
  • Flexible paid time off plan.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →