3 часа назад
AI Systems, Training
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Systems, Training (Distributed ML Systems): Building a next-generation ML model training platform for generative vision, language, and world models with an accent on distributed training, hardware-software co-design, and low-level kernel optimization. Focus on scaling multi-node systems, implementing elastic sharding and resilient checkpointing, benchmarking MFU and memory bandwidth, and translating model requirements into infrastructure and hardware specifications.
Location: US Remote; company office in Mountain View, California, with complimentary meals available at the Palo Alto office.
Company
is developing energy-efficient computing foundations for AI by mapping neural networks more directly to semiconductor device physics.
What you will do
- Build and maintain optimized, model-specific training stacks for generative vision, language, and world models.
- Design and scale multi-node distributed training systems with elastic sharding and high-throughput data streaming pipelines.
- Implement robust model checkpointing and recovery mechanisms.
- Develop and optimize kernels using CUDA and Triton.
- Create benchmarking suites for Model FLOPs Utilization, memory bandwidth, and convergence stability.
- Collaborate with theorists and infrastructure and hardware engineers to translate algorithmic trade-offs and model requirements into concrete system specifications.
Requirements
- MS, PhD, or equivalent research or project experience in AI/Machine Learning, Computer Science, Physics, Electrical Engineering, Applied Mathematics, or a related quantitative field.
- Deep expertise in modern ML software systems and in mapping transformer, Mixture of Experts, and diffusion model architectures to system performance.
- Strong understanding of cluster-level model partitioning, communication primitives, and parallelism strategies.
- Production experience implementing, debugging, and maintaining training frameworks such as Megatron-LM, DeepSpeed, Ray, or PyTorch Lightning.
- Ability to work remotely from the United States.
Nice to have
- Experience co-designing algorithms for computing paradigms closely aligned with underlying system physics.
Culture & Benefits
- Opportunity to contribute to foundational AI computing technology focused on reducing energy constraints.
- High ownership and responsibility as a foundational team member.
- Comprehensive health benefits.
- 401(k) matching.
- Unlimited paid time off and complimentary meals at the Palo Alto office.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →