Назад
Company hidden
11 дней назад

Technical Lead, Multimodal Research (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Lead, Multimodal Research (AI): Building multimodal video-understanding infrastructure that indexes physical-AI datasets across video, lidar, radar, and simulation data with an accent on model selection, evaluation, and production inference at petabyte scale. Focus on designing cost-efficient model cascades and routing, defining evaluation standards, and translating customer research needs into scalable datasets and technical programs.

Location: San Francisco, United States; 4 days/week in the SF Mission District office

Company

hirify.global builds open-source and commercial infrastructure for multimodal AI, enabling Physical AI companies to search large-scale video and sensor corpora and create training datasets or actionable alerts.

What you will do

  • Own the modeling strategy across the platform, including model families, representations, training approaches, and research priorities.
  • Move computer vision and multimodal approaches from prototypes to production inference across hundreds of thousands of video hours and petabyte-scale corpora.
  • Define model benchmarks, evaluation standards, and customer acceptance criteria.
  • Optimize the cost and performance of video understanding through distillation, cascades, routing, quantization, GPU utilization, and throughput improvements.
  • Translate customer research needs into technical programs covering taxonomies, model plans, datasets, and quality instrumentation.
  • Set the technical direction for multimodal research as a senior individual contributor without direct reports.

Requirements

  • 5+ years of experience in applied computer vision or multimodal machine learning.
  • PhD or MS in computer science, electrical engineering, robotics, or applied mathematics with a computer vision or machine learning focus, or a comparable publication or production record.
  • Strong knowledge of VLMs, VQA, embeddings, representation learning, detection, tracking, segmentation, and retrieval.
  • Hands-on experience training and evaluating models at scale on real video and sensor data, including PyTorch prototyping, inference performance, GPU utilization, throughput, and cost.
  • Experience in perception or multimodal work at a self-driving, robotics, Physical AI, frontier research, or visual-data company.

Nice to have

  • Publications at CVPR, ICCV, ECCV, NeurIPS, ICML, or ICLR.
  • Experience building or fine-tuning VLMs or multimodal foundation models.
  • Experience with long-form video, temporal reasoning, embeddings, retrieval, or content-aware indexing at scale.
  • Experience with lidar, radar, depth, simulation output, evaluation frameworks, labeling taxonomies, annotation programs, or large-scale GPU optimization.

Culture & Benefits

  • In-person, tight-knit team working 4 days per week from the San Francisco office.
  • Competitive compensation and startup equity.
  • Health, vision, dental, and 401(k) coverage with company match.
  • Flexible PTO, commuter benefit, and catered lunches and dinners for San Francisco employees.
  • Latest Apple equipment, team-building events, and poker nights.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →