Назад
Company hidden
5 дней назад

Computer Vision Applied Research Scientist (AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Computer Vision Applied Research Scientist (AI): Building and shipping a multimodal foundation model that understands construction drawings, with an accent on vision architecture, self-supervised pretraining, supervised fine-tuning, and production inference. Focus on designing multi-stage perception and relational reasoning systems, running rigorous experiments on real drawings, and improving model accuracy across construction trades and scopes.

Location: United States (Remote); collaboration during California business hours

Company

hirify.global is an early-stage AI platform for construction that embeds AI agents into estimating and bid-management workflows.

What you will do

  • Design and evaluate novel multimodal vision architectures for construction drawing understanding, including perception, text-object association, and relational reasoning.
  • Drive decisions on backbones, decoders, fusion strategies, loss functions, and training regimes.
  • Run rigorous baselines, ablations, and held-out evaluations on real construction drawings.
  • Own supervised training and self-supervised pretraining with PyTorch and modern computer vision stacks such as YOLO, SAM, and DINO.
  • Move successful models from research notebooks into production inference pipelines, collaborating on deployment, quantization, and serving.
  • Define evaluation datasets and metrics, investigate real customer failure modes, and communicate findings to engineering leadership.

Requirements

  • Must be based in the United States.
  • 7+ years of computer vision research experience or equivalent experience in an industry research lab, applied science team, or PhD research and industry.
  • Deep hands-on experience with multimodal vision transformers and dense prediction tasks such as segmentation, detection, or joint vision-language tasks.
  • Production experience with modern vision transformer backbones, including SAM, DINOv2/v3, CLIP, SigLIP, or similar models.
  • Strong PyTorch fluency and experience training large vision models, with deep fundamentals in optimization, loss design, regularization, and self-supervised learning.
  • Clear written and verbal English communication and availability during California business hours.

Nice to have

  • Graph Neural Networks or relational reasoning architectures.
  • Text spotting, OCR, or scene text detection integrated with vision models.
  • LoRA, adapters, or parameter-efficient fine-tuning of large vision models.
  • Experience with engineering drawings, document understanding, or layout analysis.
  • Open-source contributions or publications at CVPR, ICCV, ECCV, NeurIPS, or ICLR.

Culture & Benefits

  • High autonomy to propose, defend, and run research experiments.
  • Publication-friendly environment supporting research publication at top venues.
  • Direct ownership of a foundation model for construction drawings and its impact on real customers.
  • Collaboration with experienced engineering and research professionals from leading technology companies and institutions.
  • Meaningful equity in an early-stage, well-funded startup.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →