Назад
Company hidden
обновлено 4 дня назад

ML Infrastructure Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
France/UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Infrastructure Engineer (AI): Building the systems behind LLM post-training, RL, evaluation, inference, and agentic development workflows with an accent on high-throughput training pipelines and data control systems. Focus on designing robust infrastructure that directly affects model learning dynamics, training stability, and product quality.

Location: Hybrid in Paris or London; relocation package available for work from Paris.

Company

hirify.global is an AI Safety company building a safety, reliability, and optimization layer for AI systems through policy testing, enforcement, and continuous improvement.

What you will do

  • Build scalable reinforcement learning and LLM post-training pipelines, including smoke-tuning runs and approach ablations.
  • Design data control systems for rollouts, replay, filtering, evaluation, and policy updates.
  • Optimize training and inference across networking, memory, compute scheduling, storage, checkpointing, and I/O.
  • Investigate how infrastructure choices affect learning dynamics, evaluation quality, model behavior, and training stability.
  • Build experiment infrastructure for runs, artifacts, evaluations, dashboards, failure inspection, reproducibility, and cost visibility.
  • Develop agentic environments with coding-agent harnesses, browser and tool integrations, runtime sandboxes, repository-aware workflows, and multi-agent orchestration.

Requirements

  • Experience designing, building, or maintaining distributed reinforcement learning or post-training systems at scale.
  • Familiarity with PyTorch or JAX and proficiency in Python, including concurrency, asynchronous programming, multiprocessing, and performance optimization.
  • Ability to debug distributed GPU workloads across CUDA, containers, drivers, NCCL or equivalent communication layers, networking, storage, scheduling, and checkpointing.
  • Experience with profiling tools such as PyTorch Profiler, Nsight, perf, tracing, metrics, logs, or custom instrumentation.
  • Experience with inference stacks such as vLLM, SGLang, TensorRT-LLM, Dynamo, or custom serving infrastructure.
  • Strong ownership and ability to turn ambiguous infrastructure problems into working systems and improve them through feedback.

Nice to have

  • Open-source contributions, benchmarks, papers with code, or technical writing related to RL, distributed ML, LLMs, inference, evaluation, or agent infrastructure.
  • Experience in high-bar AI infrastructure, research, or model environments.
  • Experience owning custom training frameworks, fine-tuning pipelines, trainers, schedulers, checkpointing, data loaders, or performance tooling.
  • Experience with agentic coding systems, GPU clusters, Kubernetes, Slurm, Ray, custom schedulers, or cloud GPU orchestration.
  • Knowledge of NCCL, UCX, NVSHMEM, RDMA, InfiniBand, RoCE, EFA, Rust, C++, CUDA, or Go.

Culture & Benefits

  • Small, focused team working on AI safety and complex infrastructure problems.
  • Competitive compensation package including equity.
  • Flexible time off and paid time off aligned with local regulations.
  • Medical insurance in France, learning and development support, and necessary hardware, tools, and services.
  • Covered subscriptions for AI agents and IDEs, plus team off-sites twice a year.

Hiring process

  • 25-minute introductory call with HR.
  • Take-home test assignment.
  • Technical interview with the Head of Applied Research followed by a 45-minute final conversation with the CEO.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →