Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Research Infrastructure Engineer (AI): Building distributed training infrastructure, experiment orchestration systems, data pipelines, and tooling for AI research at thousands-of-GPUs scale with an accent on reliability, performance optimization, and large-scale parallelism. Focus on debugging non-deterministic training failures, improving throughput and memory efficiency, and scaling coding-agent rollouts across VM sandboxes.
Location: San Francisco, United States; on-site
Company
Windsurf builds Devin, an AI software engineer, and develops systems for advanced AI research and deployment.
What you will do
- Build and operate distributed training infrastructure for large-scale GPU clusters, including job launchers, checkpointing, recovery, fault tolerance, and monitoring.
- Scale coding-agent rollouts across VM sandboxes and support reinforcement learning workloads with hundreds of thousands of concurrent executions.
- Profile and optimize training throughput, data loading, communication, memory utilization, and compute efficiency.
- Design experiment orchestration, tracking, analysis, and research tooling that reduces iteration time.
- Build reliable, high-throughput data pipelines for training and evaluation with strong reproducibility and data quality.
- Diagnose failures across GPUs, networking, numerics, and data while implementing parallelism strategies and graceful recovery.
Requirements
- Deep experience building and operating distributed training systems for large models, from cluster infrastructure through the training loop.
- Strong systems engineering fundamentals across distributed systems, networking, storage, and the hardware-software stack.
- Proficiency in Python and C++, with systems-level experience using PyTorch or equivalent deep learning frameworks.
- Hands-on experience with GPU profiling, memory optimization, compute efficiency, and data, tensor, pipeline, or sequence parallelism.
- Strong debugging skills for complex, non-deterministic distributed systems and sufficient machine learning knowledge to work directly with researchers.
- Demonstrated capability is valued more than credentials; a PhD is only one possible signal.
Culture & Benefits
- Small, highly selective team where research and product development move together.
- Prototypes can reach real deployment quickly.
- Infrastructure operates across thousands of GPUs with access to the systems required for the work.
- Environment emphasizes speed, autonomy, technical depth, and minimal process overhead.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Senior / Staff ML Ops Engineer (AI)
184 000 - 272 000$
Microsoft AI
6 дней назад
Software Engineer (AI Infra and Model Foundry)
142 800 - 274 800$
5 дней назад
Staff Engineer, Test Automation (AI)
7 дней назад
ML Infrastructure Engineer (AI)
10 дней назад
Software Engineer, ML Infrastructure, Level 5 (AI)
178 000 - 313 000$
11 дней назад