6 дней назад
Applied ML Engineer (Vision-Language Models)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Applied ML Engineer (Vision-Language Models): Own and deliver applied post-training work for vision-language models end-to-end, including data generation, fine-tuning, evaluation, and customer engagement with an accent on visual data quality, multimodal evaluation, and real-world deployment. Focus on designing and executing multimodal post-training workflows, running supervised fine-tuning and reinforcement learning, and shaping applied multimodal AI at a foundation model company.
Location
Location: San Francisco, Boston, or Remote (Hybrid)
Company
Spun out of MIT CSAIL, builds general-purpose AI systems optimized for deployment across data centers and on-device hardware, partnering with enterprises in consumer electronics, automotive, life sciences, and financial services.
What you will do
- Own enterprise customer vision-language model (VLM) post-training projects end-to-end, from requirements to delivery and evaluation.
- Translate customer needs into multimodal post-training specifications and workflows.
- Design and execute visual data generation, filtering, annotation pipelines, and synthetic data generation for visual tasks.
- Run supervised fine-tuning, preference alignment, and reinforcement learning workflows for VLMs.
- Design task-specific evaluations for visual understanding, grounding, OCR, document parsing, and other multimodal capabilities.
- Feed evaluation results back into core post-training pipelines and contribute to baseline model development.
Requirements
- Hands-on experience with data generation and evaluation for VLM or multimodal post-training
- Experience training or fine-tuning vision-language models using supervised fine-tuning, preference alignment, and/or reinforcement learning
- Strong intuition for visual data quality, annotation design, and multimodal evaluation
- Familiarity with vision encoders, image-text architectures, and interaction with language model backbones
- English: C1 level or higher required
Nice to have
- Experience with visual grounding, document understanding, OCR, or video understanding tasks.
- Experience contributing to shared or general-purpose multimodal post-training infrastructure.
- Prior exposure to customer-facing or applied ML delivery environments.
- Familiarity with alignment or reinforcement learning techniques beyond basic supervised fine-tuning in multimodal settings.
Culture & Benefits
- Real ML work involving fine-tuning vision-language models and generating multimodal data feeding directly into core model development.
- Competitive base salary with equity in a unicorn-stage startup.
- 100% medical, dental, and vision premiums paid for employees and dependents.
- 401(k) matching up to 4% of base pay.
- Unlimited PTO plus company-wide Refill Days throughout the year.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Research Engineer (Reinforcement Learning)
17 часов назад
AI Engineer (ML/LLM)
224 000 - 344 000$
6 дней назад
Customer Engineer (ML/AI)
170 000 - 199 000$
3 дня назад
Applied AI Engineer
200 000 - 350 000$
2 дня назад
AI Research Engineer (AI)
6 дней назад
Staff AI/ML Engineer (LLMs)
250 000 - 350 000$