Назад
Company hidden
6 дней назад

Applied ML Engineer (Vision-Language Models)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Applied ML Engineer (Vision-Language Models): Own and deliver applied post-training work for vision-language models end-to-end, including data generation, fine-tuning, evaluation, and customer engagement with an accent on visual data quality, multimodal evaluation, and real-world deployment. Focus on designing and executing multimodal post-training workflows, running supervised fine-tuning and reinforcement learning, and shaping applied multimodal AI at a foundation model company.

Location

Location: San Francisco, Boston, or Remote (Hybrid)

Company

Spun out of MIT CSAIL, hirify.global builds general-purpose AI systems optimized for deployment across data centers and on-device hardware, partnering with enterprises in consumer electronics, automotive, life sciences, and financial services.

What you will do

  • Own enterprise customer vision-language model (VLM) post-training projects end-to-end, from requirements to delivery and evaluation.
  • Translate customer needs into multimodal post-training specifications and workflows.
  • Design and execute visual data generation, filtering, annotation pipelines, and synthetic data generation for visual tasks.
  • Run supervised fine-tuning, preference alignment, and reinforcement learning workflows for VLMs.
  • Design task-specific evaluations for visual understanding, grounding, OCR, document parsing, and other multimodal capabilities.
  • Feed evaluation results back into core post-training pipelines and contribute to baseline model development.

Requirements

  • Hands-on experience with data generation and evaluation for VLM or multimodal post-training
  • Experience training or fine-tuning vision-language models using supervised fine-tuning, preference alignment, and/or reinforcement learning
  • Strong intuition for visual data quality, annotation design, and multimodal evaluation
  • Familiarity with vision encoders, image-text architectures, and interaction with language model backbones
  • English: C1 level or higher required

Nice to have

  • Experience with visual grounding, document understanding, OCR, or video understanding tasks.
  • Experience contributing to shared or general-purpose multimodal post-training infrastructure.
  • Prior exposure to customer-facing or applied ML delivery environments.
  • Familiarity with alignment or reinforcement learning techniques beyond basic supervised fine-tuning in multimodal settings.

Culture & Benefits

  • Real ML work involving fine-tuning vision-language models and generating multimodal data feeding directly into core model development.
  • Competitive base salary with equity in a unicorn-stage startup.
  • 100% medical, dental, and vision premiums paid for employees and dependents.
  • 401(k) matching up to 4% of base pay.
  • Unlimited PTO plus company-wide Refill Days throughout the year.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →