Назад
Company hidden
3 часа назад

AI Systems, Training

Формат работы
remote (только USA)
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Systems, Training (Distributed ML Systems): Building a next-generation ML model training platform for generative vision, language, and world models with an accent on distributed training, hardware-software co-design, and low-level kernel optimization. Focus on scaling multi-node systems, implementing elastic sharding and resilient checkpointing, benchmarking MFU and memory bandwidth, and translating model requirements into infrastructure and hardware specifications.

Location: US Remote; company office in Mountain View, California, with complimentary meals available at the Palo Alto office.

Company

hirify.global is developing energy-efficient computing foundations for AI by mapping neural networks more directly to semiconductor device physics.

What you will do

  • Build and maintain optimized, model-specific training stacks for generative vision, language, and world models.
  • Design and scale multi-node distributed training systems with elastic sharding and high-throughput data streaming pipelines.
  • Implement robust model checkpointing and recovery mechanisms.
  • Develop and optimize kernels using CUDA and Triton.
  • Create benchmarking suites for Model FLOPs Utilization, memory bandwidth, and convergence stability.
  • Collaborate with theorists and infrastructure and hardware engineers to translate algorithmic trade-offs and model requirements into concrete system specifications.

Requirements

  • MS, PhD, or equivalent research or project experience in AI/Machine Learning, Computer Science, Physics, Electrical Engineering, Applied Mathematics, or a related quantitative field.
  • Deep expertise in modern ML software systems and in mapping transformer, Mixture of Experts, and diffusion model architectures to system performance.
  • Strong understanding of cluster-level model partitioning, communication primitives, and parallelism strategies.
  • Production experience implementing, debugging, and maintaining training frameworks such as Megatron-LM, DeepSpeed, Ray, or PyTorch Lightning.
  • Ability to work remotely from the United States.

Nice to have

  • Experience co-designing algorithms for hirify.global computing paradigms closely aligned with underlying system physics.

Culture & Benefits

  • Opportunity to contribute to foundational AI computing technology focused on reducing energy constraints.
  • High ownership and responsibility as a foundational team member.
  • Comprehensive health benefits.
  • 401(k) matching.
  • Unlimited paid time off and complimentary meals at the Palo Alto office.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →