2 часа назад
Member of Technical Staff, Infrastructure & Training Systems (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, Infrastructure & Training Systems (AI): Building distributed training infrastructure, reusable frameworks, and performance tooling for large-scale biological world models with an accent on systems performance, reliability, and hardware efficiency. Focus on optimizing communication, memory, kernels, compilation paths, fault tolerance, observability, and reproducible experimentation across evolving multimodal and long-context architectures.
Location: On-site in San Francisco or Tokyo. U.S. work authorization is required for employment in the United States.
Company
is an AI research lab developing generative genomics and biological world models to advance scientific understanding, medical discovery, and biological safety.
What you will do
- Design and scale distributed training infrastructure for large-scale biological world models.
- Optimize communication patterns, memory efficiency, custom kernels, compilation paths, and systems instrumentation.
- Build reusable internal libraries, abstractions, and workflows for reproducible and reliable model training.
- Improve fault tolerance, checkpointing, monitoring, debugging, experiment hygiene, and incident analysis.
- Collaborate with model researchers, training scientists, and data and infrastructure engineers to remove bottlenecks.
- Adapt infrastructure for multimodal models, long-context training, and evolving model architectures.
Requirements
- Strong engineering experience in distributed systems, high-performance ML infrastructure, training systems, or a related field.
- Proficiency with Python, PyTorch, Triton, CUDA, and C++.
- Strong understanding of modern deep learning frameworks and their systems internals.
- Ability to debug distributed training, performance regressions, memory issues, and reliability problems in large codebases.
- Excellent written and verbal communication across technical and scientific domains.
- Comfort collaborating with researchers, engineers, and domain experts with a strong bias toward initiative and execution.
Nice to have
- Experience with large-scale distributed training for frontier or foundation models.
- Open-source contributions to ML systems or infrastructure such as PyTorch, Torchtitan, or Megatron-LM.
- Familiarity with ML runtimes, compilers, numerics, communication libraries, or custom kernel development.
- Experience improving researcher productivity through infrastructure and developer tooling.
- Background in applied mathematics, systems, computational biology, or related quantitative sciences.
Culture & Benefits
- Work on distributed training, model architecture, and numerics problems for real biological applications.
- Collaborate across AI labs, biotech companies, hospital systems, and research institutes.
- Culture emphasizing rigor, creativity, and cross-disciplinary partnership.
- Competitive compensation, comprehensive benefits, and support for continual learning.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 часа назад
Member of Technical Staff, Kernels (AI)
200 000 - 350 000$
4 часа назад
Member of Technical Staff (Research Infrastructure)
200 000 - 400 000$
3 часа назад
Member of Technical Staff — Inference-Multi-Hardware (AI Infrastructure)
200 000 - 400 000$
3 часа назад
Member of Technical Staff — Developer Technology (AI)
200 000 - 400 000$
3 часа назад
Member of Technical Staff, Training Infra (AI)
200 000 - 350 000$
3 часа назад