Назад
Company hidden
5 дней назад

ML Infrastructure Tech Lead (AI)

200 000 - 300 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

ML Infrastructure Tech Lead (AI): Building and optimizing high-performance model training and inference systems for an agentic document platform with an accent on GPU utilization, distributed systems, and Kubernetes. Focus on optimizing model serving kernels, designing reliable multi-node training systems, and reducing latency and operational costs.

Location: On-site in San Francisco

Salary: $200k – $300k + Equity

Company

hirify.global is an agentic document platform providing a comprehensive toolkit for enterprise-scale document workflows using frontier AI models.

What you will do

  • Own the technical direction and roadmap for the company's ML infrastructure.
  • Build and maintain the training and inference stack, balancing experimentation with production performance.
  • Optimize model serving at every layer, including kernels, runtimes, batching, and scheduling.
  • Design systems for reliable multi-node, multi-GPU training and inference.
  • Improve GPU utilization, latency, throughput, and overall cost efficiency.
  • Build tooling and abstractions to help ML engineers move quickly from experiments to production.

Requirements

  • 5+ years of experience building production infrastructure, with significant focus on ML systems.
  • Proven track record of leading complex technical projects from ambiguous problems to production.
  • Strong Python and systems-engineering skills.
  • Deep understanding of modern GPU training and inference workload performance characteristics.
  • Proficiency with Kubernetes and distributed training or serving frameworks.
  • Must be based in or able to work on-site in San Francisco.

Nice to have

  • Experience optimizing or implementing CUDA, Triton, or custom model-serving kernels.
  • Contributions to frameworks such as vLLM, SGLang, PyTorch, TensorRT-LLM, or Ray.
  • Experience operating distributed inference or training across hundreds or thousands of GPUs.
  • Experience working at an early-stage or high-growth startup.

Culture & Benefits

  • Unlimited PTO for recharging.
  • Daily free lunch provided in the office.
  • Commuter reimbursement for transportation costs.
  • Comprehensive medical, dental, and vision insurance.
  • Health and wellness budget of up to $150 per month.
  • Flexible parental leave.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →