Назад
Company hidden
2 часа назад

Member of Technical Staff, Infrastructure & Training Systems (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US/Japan
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Infrastructure & Training Systems (AI): Building distributed training infrastructure, reusable frameworks, and performance tooling for large-scale biological world models with an accent on systems performance, reliability, and hardware efficiency. Focus on optimizing communication, memory, kernels, compilation paths, fault tolerance, observability, and reproducible experimentation across evolving multimodal and long-context architectures.

Location: On-site in San Francisco or Tokyo. U.S. work authorization is required for employment in the United States.

Company

hirify.global is an AI research lab developing generative genomics and biological world models to advance scientific understanding, medical discovery, and biological safety.

What you will do

  • Design and scale distributed training infrastructure for large-scale biological world models.
  • Optimize communication patterns, memory efficiency, custom kernels, compilation paths, and systems instrumentation.
  • Build reusable internal libraries, abstractions, and workflows for reproducible and reliable model training.
  • Improve fault tolerance, checkpointing, monitoring, debugging, experiment hygiene, and incident analysis.
  • Collaborate with model researchers, training scientists, and data and infrastructure engineers to remove bottlenecks.
  • Adapt infrastructure for multimodal models, long-context training, and evolving model architectures.

Requirements

  • Strong engineering experience in distributed systems, high-performance ML infrastructure, training systems, or a related field.
  • Proficiency with Python, PyTorch, Triton, CUDA, and C++.
  • Strong understanding of modern deep learning frameworks and their systems internals.
  • Ability to debug distributed training, performance regressions, memory issues, and reliability problems in large codebases.
  • Excellent written and verbal communication across technical and scientific domains.
  • Comfort collaborating with researchers, engineers, and domain experts with a strong bias toward initiative and execution.

Nice to have

  • Experience with large-scale distributed training for frontier or foundation models.
  • Open-source contributions to ML systems or infrastructure such as PyTorch, Torchtitan, or Megatron-LM.
  • Familiarity with ML runtimes, compilers, numerics, communication libraries, or custom kernel development.
  • Experience improving researcher productivity through infrastructure and developer tooling.
  • Background in applied mathematics, systems, computational biology, or related quantitative sciences.

Culture & Benefits

  • Work on distributed training, model architecture, and numerics problems for real biological applications.
  • Collaborate across AI labs, biotech companies, hospital systems, and research institutes.
  • Culture emphasizing rigor, creativity, and cross-disciplinary partnership.
  • Competitive compensation, comprehensive benefits, and support for continual learning.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →