Назад
Company hidden
1 день назад

Research Engineer (ML Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer (ML Infrastructure) (Physical AI/Robotics): Building scalable infrastructure for large-scale ML model training and production inference in industrial robotics with an accent on distributed GPU training, parallelization, and efficient compute utilization. Focus on optimizing attention, kernels, and memory, building experiment-tracking systems, and debugging bottlenecks across the training stack.

Location: Palo Alto, United States; on-site

Company

hirify.global develops generalized physical AI and robotic systems for adaptive, reasoning-intensive work in real-world industrial environments.

What you will do

  • Design and implement scalable systems for training large ML models.
  • Develop distributed training systems across hundreds of GPUs, including parallelization, sharding, and compute utilization strategies.
  • Improve training efficiency through attention optimization, kernel fusion, and memory management.
  • Build tools for experiment tracking, monitoring, debugging, and researcher productivity.
  • Debug training bottlenecks and track failures, performance, and resource utilization.
  • Support model inference and cloud infrastructure for training workloads, including efficient compute resource management.

Requirements

  • Strong experience building infrastructure for large-scale ML training.
  • Deep understanding of training and scaling modern LLM and VLM systems.
  • Proven experience setting up and scaling distributed training across hundreds of GPUs.
  • Strong understanding of data, model, and pipeline parallelism.
  • Strong Python programming skills.
  • Expert-level proficiency in PyTorch and/or JAX, plus knowledge of attention optimization, kernel fusion, and efficient memory usage.

Nice to have

  • Experience supporting inference systems in production.
  • Familiarity with robotics or embodied AI workloads.
  • Experience building experiment management tools for researchers.

Culture & Benefits

  • Hands-on work with physical AI and robotics in industrial environments.
  • Ownership-focused environment centered on solving difficult engineering problems.
  • Close collaboration with modeling teams to accelerate iteration and reduce training costs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →