Назад
Company hidden
5 часов назад

Research Scientist, Video Understanding

200 000 - 250 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Scientist, Video Understanding (Video Representation Learning and Robotics): Training large-scale video representation and video-language models on egocentric and stereo robotics data with an accent on temporal representation learning, multimodal pretraining, and rigorous evaluation. Focus on designing model architectures, running distributed PyTorch experiments, and turning trained checkpoints into reliable production embeddings and signals.

Location: New York, on-site

Salary: $200K–$250K annually, plus equity

Company

hirify.global builds data infrastructure and high-quality real-world datasets for robotics and embodied AI.

What you will do

  • Own the video understanding research agenda across manipulation, locomotion, daily activity, and long-horizon behavior.
  • Design model architectures and training strategies for large-scale video representation and video-language models.
  • Run self-supervised and multimodal pretraining with rigorous evaluations, clean ablations, and regression tracking.
  • Train and fine-tune video encoders, temporal transformers, joint-embedding models, and multimodal models.
  • Incorporate pose, depth, camera motion, and optical flow priors when they improve representation quality.
  • Convert model checkpoints into production-ready embeddings and outputs for retrieval, labeling, QA, and analytics.

Requirements

  • Deep experience training large models in PyTorch or an equivalent framework, including multi-GPU or distributed training.
  • Strong understanding of modern video representation learning and/or multimodal modeling.
  • Experience with large-scale video encoders, video-language models, VLMs, VLAs, or temporal representation learning.
  • Ability to run rigorous experiments and communicate results clearly.
  • Strong software engineering discipline and ability to write research code that can be shipped.
  • Apply to only one Research Scientist position; additional applications may result in rejection of all research submissions.

Nice to have

  • Experience with video VLM or VLA-adjacent systems such as VideoCLIP, InstructBLIP-Video, or LLaVA-Video-class models.
  • Experience with egocentric or embodied datasets such as Ego4D, EgoExo4D, EPIC-Kitchens, or Something-Something.

Culture & Benefits

  • High standards for technical excellence and first-principles reasoning.
  • Truth-seeking culture with rigorous measurement and experimentation.
  • High-intensity environment with rapid execution and end-to-end ownership.
  • Work on scarce egocentric embodied video data and research that directly feeds production systems.
  • Equity offered in addition to salary.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →