Назад
Company hidden
11 часов назад

Engineering Manager (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
lead/head
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Engineering Manager (AI): Leading the inference serving team to build and orchestrate large-scale GPU compute infrastructure with an accent on model serving, fleet orchestration, and performance optimization. Focus on hands-on architecture, scaling inference platforms to thousands of GPUs, and driving technical strategy for LLM serving engines.

Company

hirify.global is building unified general intelligence capable of generating, understanding, and operating in the physical world, with a core focus on multimodal AI.

Location: Hybrid role based in Redwood City, CA

What you will do

  • Lead, grow, and mentor the inference engineering team through hiring, coaching, and incident response.
  • Architect and build core platform components for inference serving, including routing, scheduling, and fleet-wide orchestration.
  • Set the technical roadmap for serving engines, autoscaling, caching, and observability.
  • Own platform SLOs, including latency, availability, GPU utilization, and cost per generation.
  • Partner with research teams to integrate new architectures into production and online evaluation loops.
  • Maintain a hands-on approach, spending at least 50% of time building, debugging, and solving complex design challenges.

Requirements

  • 8+ years of experience in large-scale distributed systems or ML infrastructure.
  • Proven track record of building and operating model-serving platforms at the scale of thousands of GPUs.
  • Deep expertise in LLM serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong command of continuous batching, KV-cache management, quantization, and parallelism strategies.
  • Proficiency in Python, PyTorch, and Kubernetes at scale.
  • Technical leadership experience managing teams through rapid growth.

Nice to have

  • Experience serving multimodal generative models (diffusion, video) and multimedia processing.
  • Knowledge of modern networking stacks like RDMA (RoCE, InfiniBand) and NVLink.
  • Experience with heterogeneous accelerators (AMD, TPU, Trainium).
  • Contributions to open-source infrastructure (Ray, Kubernetes, vLLM).
  • Systems-language depth in Rust, C++, or CUDA/HIP for runtime optimization.

Culture & Benefits

  • Opportunity to work on cutting-edge multimodal AI infrastructure.
  • Hands-on leadership role that prioritizes technical building over pure management.
  • Collaborative environment working closely with research and engineering teams.
  • Commitment to equal opportunity employment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →