Назад
Company hidden
13 дней назад

Software Engineer, Distributed Training (AI)

200 000 - 420 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer, Distributed Training (AI): Building and optimizing distributed training engines for large language models and reinforcement-learning workloads with an accent on GPU efficiency, numerical correctness, and reliable execution across large clusters. Focus on coordinating rollouts and weight transfers, implementing checkpoint recovery, and diagnosing complex distributed and concurrent failures.

Location: Palo Alto, California, United States

Salary: $200,000–$420,000 USD per year

Company

hirify.global is building personal AI through local-inference hardware, bespoke training infrastructure, next-generation user interfaces, and deep-learning research.

What you will do

  • Build and optimize distributed training systems for dense and mixture-of-experts models, including low-rank adapter training.
  • Improve reinforcement-learning pipelines by coordinating sampling, reward computation, training updates, and weight transfer.
  • Optimize GPU memory usage, parallelism, communication, and training throughput.
  • Implement checkpointing, resumption, and worker recovery while preserving consistent training state.
  • Validate losses, gradients, and optimizer behavior and diagnose numerical or distributed execution failures.
  • Partner with researchers to implement new learning algorithms and expose them through the River API.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Hands-on experience building or substantially improving distributed model-training systems.
  • Strong proficiency in Python and a modern deep-learning framework such as PyTorch or JAX.
  • Understanding of backpropagation, optimizers, mixed-precision training, and GPU memory management.
  • Strong debugging skills across concurrent execution, collective communication, and distributed failure recovery.
  • Collaborative mindset and willingness to work across the stack.

Nice to have

  • Experience with reinforcement-learning infrastructure, rollout generation, or asynchronous training.
  • Familiarity with tensor, pipeline, expert, or data parallelism and their performance tradeoffs.
  • Experience with LoRA, mixture-of-experts training, distributed optimizers, activation checkpointing, NCCL, or communication profiling.
  • Proficiency in C++, Rust, or CUDA and experience investigating performance below the framework layer.
  • Contributions to training frameworks or experience operating large training runs.

Culture & Benefits

  • Work alongside scientists, engineers, and builders with experience scaling consumer systems and frontier-model infrastructure.
  • Comprehensive health, dental, and vision insurance.
  • Unlimited paid time off.
  • Relocation assistance is available as needed.
  • Visa sponsorship is available for qualified candidates.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →