Назад
Company hidden
4 часа назад

Applied AI Researcher (Agent Systems & Evaluation)

193 930 - 352 290$
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Applied AI Researcher (Agent Systems & Evaluation) (LLM agents and evaluation): Building closed-loop evaluation systems for production AI agents and post-training open-weight vision-language models on proprietary autonomous-driving data with an accent on real-world measurement, test-time scaling, and rigorous experimentation. Focus on constructing trustworthy eval pipelines, automated hill climbing, and quantifying the impact of agent improvements on engineering and driving workflows.

Location: Mountain View, California, United States

Salary: $193,930–$352,290 annual base pay, plus annual performance bonus, equity, and benefits.

Company

hirify.global is a self-driving technology company developing scalable autonomous-driving systems that combine AI with automotive-grade hardware.

What you will do

  • Establish the evaluation foundation for hirify.global's production agent fleet, including data sources, task suites, noise floors, and statistical acceptance standards.
  • Convert real engineering workflows into closed-loop measurements that run on every system change.
  • Research frontier model developments and turn relevant findings into experiments against real production workloads.
  • Investigate test-time scaling, sampling and search strategies, verifier-guided selection, routing, reasoning budgets, and model escalation.
  • Build automated hill-climbing workflows for prompts, context strategies, tools, routing, reasoning budgets, and model selection.
  • Run supervised fine-tuning and reinforcement-learning experiments on open-weight vision-language models using proprietary autonomous-driving data.

Requirements

  • Graduate degree in computer science, machine learning, statistics, or a related field, or equivalent research experience.
  • Deep understanding of LLM pretraining, post-training, inference, and model mechanisms.
  • Experience across the full evaluation pipeline: sourcing data, constructing evaluation loops, and automating improvement.
  • Strong experimental design skills, including understanding variance, statistical power, and production experimentation constraints.
  • Hands-on post-training experience with supervised fine-tuning and reinforcement learning, plus data curation and evaluation.
  • Strong Python skills, production-systems experience, and direct experience building, evaluating, or studying LLM agent systems.

Nice to have

  • Published or applied work in agent evaluation, reasoning, test-time compute, reinforcement learning, or verification.
  • Experience with multimodal or vision-language models and large-scale data curation.
  • Experience with A/B testing, causal inference from observational data, and offline-to-online correlation.
  • Experience building evaluation harnesses, task suites, or LLM-as-judge systems.
  • Familiarity with autonomous systems, safety cases, or verification-gated deployment.

Culture & Benefits

  • Small startup-style team operating within an established self-driving company.
  • Direct collaboration with engineering leadership and the CEO, with ownership of technical decisions.
  • Access to production agent systems, compute, proprietary driving data, and a labeling workforce.
  • Annual performance bonus, equity, and a competitive benefits package.
  • Commitment to inclusion, diversity, and psychological safety.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →