Назад
обновлено 5 дней назад

Research Infrastructure Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Infrastructure Engineer (AI): Building distributed training infrastructure, experiment orchestration systems, data pipelines, and tooling for AI research at thousands-of-GPUs scale with an accent on reliability, performance optimization, and large-scale parallelism. Focus on debugging non-deterministic training failures, improving throughput and memory efficiency, and scaling coding-agent rollouts across VM sandboxes.

Location: San Francisco, United States; on-site

Company

Windsurf builds Devin, an AI software engineer, and develops systems for advanced AI research and deployment.

What you will do

  • Build and operate distributed training infrastructure for large-scale GPU clusters, including job launchers, checkpointing, recovery, fault tolerance, and monitoring.
  • Scale coding-agent rollouts across VM sandboxes and support reinforcement learning workloads with hundreds of thousands of concurrent executions.
  • Profile and optimize training throughput, data loading, communication, memory utilization, and compute efficiency.
  • Design experiment orchestration, tracking, analysis, and research tooling that reduces iteration time.
  • Build reliable, high-throughput data pipelines for training and evaluation with strong reproducibility and data quality.
  • Diagnose failures across GPUs, networking, numerics, and data while implementing parallelism strategies and graceful recovery.

Requirements

  • Deep experience building and operating distributed training systems for large models, from cluster infrastructure through the training loop.
  • Strong systems engineering fundamentals across distributed systems, networking, storage, and the hardware-software stack.
  • Proficiency in Python and C++, with systems-level experience using PyTorch or equivalent deep learning frameworks.
  • Hands-on experience with GPU profiling, memory optimization, compute efficiency, and data, tensor, pipeline, or sequence parallelism.
  • Strong debugging skills for complex, non-deterministic distributed systems and sufficient machine learning knowledge to work directly with researchers.
  • Demonstrated capability is valued more than credentials; a PhD is only one possible signal.

Culture & Benefits

  • Small, highly selective team where research and product development move together.
  • Prototypes can reach real deployment quickly.
  • Infrastructure operates across thousands of GPUs with access to the systems required for the work.
  • Environment emphasizes speed, autonomy, technical depth, and minimal process overhead.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →