Назад
Company hidden
7 дней назад

Research Scientist, Video & Multimodal

160 000 - 185 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Scientist, Video & Multimodal (AI): Designing data specifications and evaluation methodologies for video understanding, video generation, and multimodal models with an accent on temporal reasoning, grounding, fidelity, and physical plausibility. Focus on validating data choices through fine-tuning and ablation experiments, developing adversarial benchmarks, and publishing reproducible research.

Location: Remote - United States

Salary: $160,000–$185,000 per year, based on experience, skills, and qualifications.

Company

hirify.global is a global data engineering company that provides data, evaluation frameworks, platforms, and human expertise for generative AI builders and adopters.

What you will do

  • Define data specifications, annotation schemas, sampling strategies, and evaluation criteria for video and multimodal models.
  • Develop evaluation methodologies for video understanding, including temporal grounding, long-context reasoning, event localization, retrieval, and cross-modal evaluation.
  • Develop evaluation methodologies for video generation focused on fidelity, temporal coherence, physical plausibility, and human judgment.
  • Run fine-tuning, evaluation, and ablation experiments to measure the impact of data decisions.
  • Design adversarial and stumping evaluations, convert model failures into improved data, and publish benchmarks, methodologies, and research papers.
  • Collaborate with annotation teams, subject-matter experts, customers, frontier labs, and synthetic-data specialists to operationalize collection and labeling plans.

Requirements

  • Approximately 5+ years of hands-on industry experience in video understanding or multimodal machine learning; a current PhD research agenda may offset experience at the lower end.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an MS or PhD is preferred.
  • Hands-on experience training and evaluating video or multimodal models, with strong PyTorch fundamentals.
  • Experience with ffmpeg, decord, temporal and COCO-style annotations, WebDataset, Parquet, Arrow, and HuggingFace datasets.
  • Experience fine-tuning large video or vision-language models using HuggingFace Transformers, PEFT, and efficient inference, including work with long-form video, streaming, temporal segmentation, or synthetic video generation.
  • First-author publications or strong open-source contributions at venues such as CVPR, ICCV, ECCV, NeurIPS, or ICLR, plus the ability to explain technical decisions clearly and conduct reproducible experiments.

Nice to have

  • Interest or hands-on experience in responsible-AI evaluation, red-teaming, safety testing, or robustness testing for video and multimodal systems.

Culture & Benefits

  • Direct collaboration with customers and frontier labs developing video understanding, video-language, and video-generation models.
  • Practical research focused on data quality, evaluation science, benchmarks, and measurable model improvements.
  • Opportunity to publish research that advances video and multimodal AI.
  • Remote work from the United States.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →