Назад
Company hidden
10 часов назад

Applied Data Science Role (AI)

Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Applied Data Science Role (AI): Building evaluation pipelines, datasets, graders, and metrics to measure and improve the quality of AI agents handling real-world tasks with an accent on statistical rigor, production systems, and multi-step agent behavior. Focus on calibrating LLM-as-a-judge systems, analyzing traces and tool calls, identifying failure causes, and validating improvements without unacceptable regressions.

Location: On-site in Palo Alto, California, United States; the listing also identifies the location type as remote.

Company

hirify.global develops AI agents for real-world tasks such as scheduling, email, browser automation, and business software workflows.

What you will do

  • Architect and maintain automated evaluation pipelines across agent capabilities and product surfaces.
  • Define success criteria, gold datasets, regression suites, and metrics for task success, tool selection, instruction adherence, factual consistency, latency, cost, and reliability.
  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and measure false positives, false negatives, variance, and grader agreement.
  • Analyze traces, tool calls, and production outcomes to identify root causes and develop failure taxonomies.
  • Run offline experiments and use production evidence to compare models, prompts, and capability implementations.
  • Build dashboards and reports, and partner with engineering to verify quality improvements without unacceptable regressions.

Requirements

  • 5+ years of experience in data science, machine learning, or analytics, focused on evaluation systems, metrics frameworks, or production quality measurement.
  • Experience designing and implementing evaluation and grading frameworks for production ML or AI products.
  • Production-quality Python and SQL skills for automated pipelines and analysis at scale.
  • Strong knowledge of evaluation methodology, dataset construction, metric selection, and statistical and experimental design.
  • Experience creating ground-truth data, including labeling guidelines, annotation quality control, ambiguity resolution, and dataset maintenance.
  • Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes.

Nice to have

  • Experience with agentic systems, multi-step task evaluation, or consumer-facing production ML.

Culture & Benefits

  • Work at the intersection of evaluation design, statistics, and production systems.
  • Collaborate with engineering, product, and leadership to make evaluation results clear and actionable.
  • Visa sponsorship is not available.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →