Назад
Company hidden
9 часов назад

Research Engineer (Physical AI)

150 000 - 250 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer (Physical AI) (Vision and Multimodal Models): Building production visual-understanding systems that make petabytes of video and sensor data queryable for Physical AI training with an accent on VLMs, VQA, embeddings, perception models, and corpus-scale inference. Focus on training and selecting deployable models, reducing annotation costs through distillation and batching, and shipping rich datasets into customer GPU workflows.

Location: San Francisco, with in-person work 4 days per week in the SF Mission District office

Salary: $150K–$250K annually, plus equity

Company

hirify.global develops Daft, an open-source distributed data engine and video-native indexing infrastructure for multimodal AI and Physical AI systems.

What you will do

  • Own the visual-understanding roadmap from model-family selection through production inference at corpus scale.
  • Train, fine-tune, and evaluate VLMs, VQA models, embedding models, and convolutional perception models on customer datasets and benchmarks.
  • Reduce per-clip annotation costs through model selection, distillation, batching, and decode pipelining.
  • Build rich, queryable datasets by designing taxonomies with researchers, measuring quality, and versioning outputs.
  • Partner with dataloading and storage teams to connect visual-understanding outputs to the index and GPU workflows.
  • Work directly with researchers at partner Physical AI labs and use their training iterations as the feedback loop.

Requirements

  • Strong familiarity with modern vision and multimodal models, including convolutional networks, VLMs, VQA, and embeddings.
  • Experience running vision or multimodal models at scale on real video and sensor data, ideally for detection, tracking, segmentation, retrieval, or captioning.
  • Perception experience from self-driving, robotics, visual-data, or equivalent research environments.
  • Experience with cloud infrastructure and large-scale data processing, including jobs running across thousands of GPU-hours of video.
  • Production-oriented research engineering skills, with a strong focus on data and infrastructure.

Nice to have

  • Experience training vision or multimodal models from scratch.
  • ML or AI research background demonstrated through papers, citations, or research-organization experience.
  • Hands-on experience with Spark, Ray, or Daft.
  • Experience with embeddings, retrieval, content-aware search, labeling taxonomies, or annotation programs.

Culture & Benefits

  • In-person, tight-knit engineering team working four days per week from the office.
  • Startup equity and competitive compensation.
  • Health, vision, and dental coverage.
  • 401(k) plan with company match, commuter benefit, and flexible PTO.
  • Catered lunches and dinners, team-building events, and the latest Apple equipment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →