Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Research Engineer (AI): Building multimodal data pipelines, ML training and evaluation infrastructure, simulation environments, and distributed reinforcement learning systems for observability agents with an accent on world models, autonomous incident response, and scalable AI infrastructure. Focus on orchestrating distributed training with Ray, implementing rigorous benchmarks, and turning research prototypes into reliable Datadog services.
Location: Paris, France; hybrid work
Company
Datadog is a global SaaS business that provides cloud observability, infrastructure monitoring, and security solutions.
What you will do
- Build and operate multimodal data pipelines, ML training and evaluation infrastructure, benchmarks, and internal tooling.
- Implement models and run large-scale experiments while profiling reliability, performance, and cost.
- Build simulation environments and replay infrastructure for autonomous agent training and evaluation.
- Orchestrate distributed training and reinforcement learning with Ray, including scheduling, scaling, and failure recovery.
- Establish automated benchmarks and regression tests for world models, agent performance, and simulation fidelity.
- Collaborate with Research Scientists, Product, and Engineering to integrate research capabilities into reliable Datadog services and contribute to publications and open-source artifacts.
Requirements
- Depth in distributed computing, reinforcement learning infrastructure, and ML systems for training and inference at scale.
- Proficiency in Python and familiarity with a systems language such as Rust, C++, or Go.
- Practical experience operating ML training and inference systems with PyTorch or JAX, including containerization, orchestration, and GPU acceleration.
- Experience with large-scale model training and fine-tuning, including SFT, RLVR, RLHF, quantization, or speculative decoding.
- Comfort with modern cloud and data infrastructure and the ability to explain technical and performance trade-offs clearly.
- Experience supporting or contributing to research publications.
Nice to have
- Experience with Ray, Slurm, or similar distributed computing frameworks.
- Software engineering experience in observability, SRE, or security.
- Experience bridging research prototypes with product applications involving foundation models, world models, or RL-trained agents.
- Hands-on GPU programming and optimization experience, including CUDA.
- Experience building production data pipelines or simulation environments for agent training.
Culture & Benefits
- Competitive benefits that vary by country of employment and employment arrangement.
- New-hire stock equity through RSUs and an employee stock purchase plan.
- Collaboration with colleagues across Datadog offices in New York City and Paris.
- Opportunities to attend and present at conferences and meetups.
- Mentor and buddy programs, employee resource groups, and an inclusive company culture.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Sr. AI Engineer (Agentic AI)
135 000 - 180 000$
8 дней назад
Member of Technical Staff, Machine Learning
Decagon
8 дней назад
Research Engineer (AI Safety)
200 000 - 400 000$
Anthropic
9 дней назад
Applied AI Research Engineer
300 000 - 400 000$
11 дней назад
R&D AI Manager
13 дней назад