2 часа назад
Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Engineer (AI): Building the execution, sandboxing, and observability platform for long-horizon reinforcement learning environments with an accent on state snapshotting, restoration, isolation, and scalable distributed execution. Focus on optimizing throughput, latency, and cost across thousands of concurrent rollouts, while enabling researchers and enterprise customers to debug and operate reliable agent environments.
Location: Hybrid in Mountain View, California, United States
Company
is an applied AI research lab building data and reinforcement learning environments for training and evaluating agents.
What you will do
- Design and own the sandboxing and execution layer for realistic, multi-tool RL environments.
- Build snapshot, restore, inspection, and branching capabilities for disk, process, memory, and accelerator state where relevant.
- Develop failure detection and recovery systems for reward hacks, infrastructure faults, and fairness issues.
- Extend execution to long-horizon and multi-node environments, optimizing throughput, latency, utilization, and cost per rollout.
- Build the framework for specifying, packaging, deploying, debugging, and observing RL environments across thousands of concurrent rollouts.
- Scale prototypes into reproducible production systems and create documentation and tools for internal and external users.
Requirements
- Production systems or research infrastructure experience at scale, including distributed systems, execution engines, or container and sandboxing infrastructure.
- Strong knowledge of containers, isolation, namespaces, cgroups, virtual machines, filesystems, process management, and state management.
- Experience with profiling, scheduling, resource utilization, cost optimization, and distributed computing.
- Proficiency with cloud platforms such as GCP or AWS.
- Strong Python skills and a systematic approach to testing, validation, and reliability.
- Excellent communication skills and the ability to translate research needs into infrastructure requirements.
Nice to have
- Experience with RL training or evaluation infrastructure and agent rollout execution layers.
- Experience with checkpointing, snapshot-restore systems, CRIU, or distributed state management.
- Rust, Go, or C++ experience.
- Background in high-throughput, low-latency execution systems or research engineering at an AI or systems-focused company.
- Contributions to infrastructure, datasets, benchmarks, or open-source systems.
Culture & Benefits
- Work closely with research and data teams.
- Collaborate directly with frontier AI labs and enterprise customers.
- Health coverage is provided.
- Compensation includes salary and equity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →