Назад
Company hidden
2 часа назад

Head of Engineering (AI Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Грейд
head
Английский
b2
Страна
US
Релокация
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Head of Engineering (AI Infrastructure): Building and leading the engineering organization developing the systems that power vLLM and high-performance AI inference with an accent on GPU optimization, inference runtimes, ML systems, and hardware-software co-design. Focus on recruiting senior ML systems talent, translating research and infrastructure work into execution plans, and delivering reliable inference across models, hardware, and deployment environments.

Location: On-site in San Francisco, California; relocation will be considered for exceptional candidates.

Company

hirify.global develops vLLM and AI inference infrastructure to make model inference faster and more cost-efficient.

What you will do

  • Build and lead the engineering organization developing systems that power vLLM and hirify.global.
  • Scale a senior-heavy, specialized engineering team in partnership with the founders.
  • Recruit, assess, develop, and retain senior engineers, staff-level ICs, PhDs, and research-oriented engineers.
  • Translate research and infrastructure initiatives into priorities, ownership, execution plans, and operating mechanisms.
  • Guide technical work across inference runtimes, GPU and accelerator optimization, kernels, memory, communication bottlenecks, and hardware-software tradeoffs.
  • Represent engineering with open-source contributors, hardware partners, cloud providers, customers, candidates, and investors.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.
  • Engineering leadership experience building and scaling specialized teams in LLM inference, ML systems, GPU or accelerator software, distributed systems, or related infrastructure.
  • Hands-on technical understanding of inference runtimes, GPU or accelerator optimization, kernels, memory and communication bottlenecks, and hardware-software tradeoffs.
  • Strong experience recruiting and retaining senior technical talent in production engineering environments.
  • Ability to set technical priorities, establish accountable ownership, and unblock teams without displacing technical ownership.

Nice to have

  • Experience with LLM serving, vLLM, SGLang, model execution, GPU kernels, compiler or runtime systems, or distributed AI infrastructure.
  • Experience integrating research-oriented or PhD talent into production engineering teams.
  • Contributions to open-source ML systems such as vLLM, SGLang, PyTorch, Ray, Triton, XLA, or ROCm.
  • Experience leading engineering in an early-stage AI infrastructure, developer infrastructure, distributed systems, or open-source company.

Culture & Benefits

  • Work alongside the creators and core maintainers of vLLM.
  • Maintain high technical standards, fast iteration, and direct ownership.
  • Competitive base compensation with meaningful equity, determined by background, skills, and experience.
  • Generous health, dental, and vision benefits.
  • 401(k) company match.
  • Visa sponsorship is available on a case-by-case basis.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →