5 часов назад
Research Scientist, Video Understanding
200 000 - 250 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Research Scientist, Video Understanding (Video Representation Learning and Robotics): Training large-scale video representation and video-language models on egocentric and stereo robotics data with an accent on temporal representation learning, multimodal pretraining, and rigorous evaluation. Focus on designing model architectures, running distributed PyTorch experiments, and turning trained checkpoints into reliable production embeddings and signals.
Location: New York, on-site
Salary: $200K–$250K annually, plus equity
Company
builds data infrastructure and high-quality real-world datasets for robotics and embodied AI.
What you will do
- Own the video understanding research agenda across manipulation, locomotion, daily activity, and long-horizon behavior.
- Design model architectures and training strategies for large-scale video representation and video-language models.
- Run self-supervised and multimodal pretraining with rigorous evaluations, clean ablations, and regression tracking.
- Train and fine-tune video encoders, temporal transformers, joint-embedding models, and multimodal models.
- Incorporate pose, depth, camera motion, and optical flow priors when they improve representation quality.
- Convert model checkpoints into production-ready embeddings and outputs for retrieval, labeling, QA, and analytics.
Requirements
- Deep experience training large models in PyTorch or an equivalent framework, including multi-GPU or distributed training.
- Strong understanding of modern video representation learning and/or multimodal modeling.
- Experience with large-scale video encoders, video-language models, VLMs, VLAs, or temporal representation learning.
- Ability to run rigorous experiments and communicate results clearly.
- Strong software engineering discipline and ability to write research code that can be shipped.
- Apply to only one Research Scientist position; additional applications may result in rejection of all research submissions.
Nice to have
- Experience with video VLM or VLA-adjacent systems such as VideoCLIP, InstructBLIP-Video, or LLaVA-Video-class models.
- Experience with egocentric or embodied datasets such as Ego4D, EgoExo4D, EPIC-Kitchens, or Something-Something.
Culture & Benefits
- High standards for technical excellence and first-principles reasoning.
- Truth-seeking culture with rigorous measurement and experimentation.
- High-intensity environment with rapid execution and end-to-end ownership.
- Work on scarce egocentric embodied video data and research that directly feeds production systems.
- Equity offered in addition to salary.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 часов назад
Machine Learning Engineer (Robotics)
100 000 - 150 000$
2 часа назад
Robot Foundation Model Researcher (AI/Robotics)
200 000 - 350 000$
4 часа назад
Staff Machine Learning Research Scientist (AI)
190 000 - 260 000$
2 часа назад
Scientist / Senior Scientist, Multimodal AI
179 400 - 330 000$
4 часа назад
Head of Computer Vision and Machine Learning (Robotics)
160 000 - 220 000$
3 часа назад
Staff ML Research Scientist (Embodied AI)
154 310 - 192 887$