12 дней назад
Research Scientist (Multimodal AI)
122 000 - 181 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Research Scientist (Multimodal AI): Developing Vision-Language Models and video foundation models for human understanding, behavior synthesis, and multimodal interaction with an accent on multimodal reasoning, temporal modeling, and generative synthesis. Focus on designing transformer-based architectures, training large-scale neural networks, evaluating models across vision, language, and video benchmarks, and analyzing failure modes.
Location: Pittsburgh, PA, United States
Salary: $122,000–$181,000 per year, plus bonus, equity, and benefits.
Company
builds social and immersive technologies that help people connect, find communities, and grow businesses.
What you will do
- Design multimodal architectures that combine vision, language, and temporal signals for human understanding.
- Develop and train Vision-Language Models for visual question answering, image-text reasoning, and grounded human-centric understanding.
- Build video foundation models for temporal reasoning, action synthesis, and long-form video synthesis.
- Research generative techniques for human-centric video, motion, and multimodal content creation.
- Run experiments across vision, language, and video benchmarks, analyze failure modes, and improve model accuracy and generalization.
- Contribute to the full research lifecycle, including problem formulation, dataset curation, model development, and evaluation.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience; the degree must be completed before joining.
- 2+ years of experience in multimodal AI research, including Vision-Language Models, video understanding, or human-centric AI systems.
- 2+ years of experience implementing and training large-scale neural networks with PyTorch and transformer-based architectures.
- Experience designing experiments and quantitatively evaluating multimodal models across vision, language, and video benchmarks.
- Production-quality or research-quality Python coding experience for multimodal AI applications.
Nice to have
- Experience developing or fine-tuning Vision-Language Models for human understanding.
- Experience with video foundation models, temporal transformers, or large-scale video pretraining.
- Published multimodal AI research at venues such as CVPR, ICCV, or NeurIPS.
- Experience with diffusion models, GANs, or autoregressive models for video or motion generation.
Culture & Benefits
- Bonus, equity, and employee benefits are provided in addition to base compensation.
- Work contributes to augmented reality, virtual reality, and the next generation of social technology.
- Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →