Назад
Company hidden
12 дней назад

Research Scientist (Multimodal AI)

122 000 - 181 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Scientist (Multimodal AI): Developing Vision-Language Models and video foundation models for human understanding, behavior synthesis, and multimodal interaction with an accent on multimodal reasoning, temporal modeling, and generative synthesis. Focus on designing transformer-based architectures, training large-scale neural networks, evaluating models across vision, language, and video benchmarks, and analyzing failure modes.

Location: Pittsburgh, PA, United States

Salary: $122,000–$181,000 per year, plus bonus, equity, and benefits.

Company

hirify.global builds social and immersive technologies that help people connect, find communities, and grow businesses.

What you will do

  • Design multimodal architectures that combine vision, language, and temporal signals for human understanding.
  • Develop and train Vision-Language Models for visual question answering, image-text reasoning, and grounded human-centric understanding.
  • Build video foundation models for temporal reasoning, action synthesis, and long-form video synthesis.
  • Research generative techniques for human-centric video, motion, and multimodal content creation.
  • Run experiments across vision, language, and video benchmarks, analyze failure modes, and improve model accuracy and generalization.
  • Contribute to the full research lifecycle, including problem formulation, dataset curation, model development, and evaluation.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience; the degree must be completed before joining.
  • 2+ years of experience in multimodal AI research, including Vision-Language Models, video understanding, or human-centric AI systems.
  • 2+ years of experience implementing and training large-scale neural networks with PyTorch and transformer-based architectures.
  • Experience designing experiments and quantitatively evaluating multimodal models across vision, language, and video benchmarks.
  • Production-quality or research-quality Python coding experience for multimodal AI applications.

Nice to have

  • Experience developing or fine-tuning Vision-Language Models for human understanding.
  • Experience with video foundation models, temporal transformers, or large-scale video pretraining.
  • Published multimodal AI research at venues such as CVPR, ICCV, or NeurIPS.
  • Experience with diffusion models, GANs, or autoregressive models for video or motion generation.

Culture & Benefits

  • Bonus, equity, and employee benefits are provided in addition to base compensation.
  • Work contributes to augmented reality, virtual reality, and the next generation of social technology.
  • Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →