2 дня назад
Senior AI Infrastructure Engineer (PyTorch)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Infrastructure Engineer (PyTorch): Building and optimizing distributed systems for training large-scale multimodal models across thousands of GPUs with an accent on advanced parallelism, training stability, and cluster utilization. Focus on implementing FSDP, Tensor Parallelism, and building robust monitoring tools for foundation-model training at scale.
Location: Redwood City, CA (Hybrid)
Company
is building unified general intelligence capable of generating, understanding, and operating in the physical world through multimodal models.
What you will do
- Design, implement, and optimize distributed training systems for models running on thousands of GPUs.
- Research and apply advanced parallelization techniques including FSDP, Tensor, Pipeline, and Expert Parallelism.
- Develop monitoring, visualization, and debugging tools to ensure large-scale training reliability.
- Optimize training stability, convergence, and resource utilization across massive GPU clusters.
Requirements
- Extensive experience with distributed PyTorch training and parallelisms in foundation-model training.
- Deep understanding of GPU clusters, networking, and storage systems.
- Proficiency with communication libraries such as NCCL and MPI.
- Proven track record of solving complex problems in distributed-system optimization.
- Must be able to work in a hybrid capacity in Redwood City, CA.
Nice to have
- Strong Linux systems administration and scripting skills.
- Experience managing training runs across 100+ GPUs.
- Familiarity with containerization, orchestration, and cloud infrastructure.
Culture & Benefits
- Opportunity to work on cutting-edge multimodal foundation models.
- Focus on high-impact infrastructure challenges at massive scale.
- Collaborative research-driven environment.
- Equal opportunity employer.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Distributed Training Infrastructure Engineer (AI)
5 дней назад
GPU Kernel Engineer (AI)
5 дней назад
Distributed LLM Inference Engineer (AI)
170 000 - 245 000$
5 дней назад
Systems/GPU Engineer (AI)
160 000 - 320 000$
5 часов назад
Senior AI Infrastructure Engineer (Fintech)
3 дня назад
Research Engineer, AI Systems
165 000 - 310 000$