Назад
Company hidden
обновлено 3 дня назад

Multimodal ML Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
France/UK
Релокация
France
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Multimodal ML Engineer (PyTorch/Multimodal AI): Training and shipping vision, audio, video, and speech models for an AI safety platform with an accent on large-scale multimodal architecture and production optimization. Focus on building alignment pipelines, optimizing MoE architectures for efficient inference, and designing evaluation metrics for complex multimodal reasoning.

Location: Hybrid work from Paris or London; relocation package available for Paris

Company

hirify.global builds an AI safety platform that tests, enforces, and continuously improves natural-language policies for AI systems.

What you will do

  • Train and fine-tune large-scale multimodal models covering vision, language, video, audio, and speech.
  • Extend models for image understanding, video temporal modeling, long-context processing, and streaming audio.
  • Design experiments involving architectures, data mixes, and training recipes.
  • Build multimodal data pipelines, including dataset curation and synthetic data generation.
  • Develop alignment pipelines using SFT, DPO, GRPO, and reward modeling across modalities.
  • Optimize and deploy models for production through quantization, distillation, batching, streaming, evaluation, and low-latency serving.

Requirements

  • 3+ years of experience training large-scale deep learning models in multimodal domains.
  • Strong PyTorch skills and hands-on distributed training experience with DeepSpeed, FSDP, or similar tools.
  • Deep understanding of multimodal architectures, including vision and audio encoders, projectors, and LLMs.
  • Hands-on experience with multimodal RLHF and alignment, including GRPO, DPO, and reward modeling.
  • Experience with video or audio sequence modeling, temporal modeling, long-context processing, efficient attention, or streaming inference.
  • Track record of shipping production models, plus strong engineering fundamentals in clean code, testing, version control, and documentation.

Nice to have

  • Understanding of audio signal processing, including spectrograms, mel features, and noise reduction.
  • Experience with MoE architectures and large-scale multimodal dataset curation.

Culture & Benefits

  • Competitive compensation package with equity.
  • Flexible time off and paid time off aligned with local regulations.
  • Hybrid work from Paris or London, with a relocation package for Paris.
  • Medical insurance in France, learning and development support, and required hardware, tools, and services.
  • Covered subscriptions for AI agents and IDEs, plus team off-sites twice a year.

Hiring process

  • 25-minute introductory call with HR.
  • Take-home test assignment.
  • Technical interview with the Head of Applied Research followed by a 45-minute final conversation with the CEO.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →