Назад
Company hidden
6 дней назад

Benchmark Engineer (AI)

250 000 - 290 000$
Формат работы
hybrid/onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Benchmark Engineer (AI): Building public, automated benchmarks and weekly leaderboards that evaluate data providers for accuracy, coverage, freshness, speed, and cost, with an accent on verified ground truth, reproducible harnesses, and defensible methodology. Focus on designing datasets and scoring systems, integrating messy vendor APIs, and feeding benchmark results back into Alexandria's provider selection.

Location: San Francisco, CA or Toronto, ON. The role is hybrid with office attendance at least 3 days per week; another description states on-site work in San Francisco five days per week. Candidates must be authorized to work in the United States or Canada.

Salary: $250,000–$290,000 USD/year in San Francisco or CA$210,000–CA$224,000 CAD/year in Toronto, plus competitive equity.

Company

hirify.global provides an API that converts web pages into clean, LLM-ready markdown and structured data for AI agents.

What you will do

  • Design and run public, head-to-head benchmarks of data providers against verified ground truth.
  • Build and maintain test datasets, labels, ground truth, scoring methods, and safeguards against stale or leaked data.
  • Own a reproducible, versioned automated benchmark harness and weekly releases.
  • Measure provider accuracy, coverage, freshness, speed, and cost across data categories.
  • Work with marketing on public leaderboard pages and communicate findings to engineers, marketers, and vendors.
  • Feed benchmark results into Alexandria to improve provider selection and expand benchmark coverage.

Requirements

  • 4+ years of experience in machine learning, research engineering, or data engineering.
  • Shipped evaluations or benchmarks and can explain why the results are trustworthy.
  • Strong Python and API integration skills, including experience with rate limits, inconsistent schemas, and messy vendor APIs.
  • Experience building test sets and scoring methods, including sampling, labeling, inter-rater agreement, and appropriate use of LLM judges.
  • Ability to work independently, ship quickly, and communicate findings clearly.
  • Authorization to work in the United States or Canada is required. US visa sponsorship is unavailable; Canadian sponsorship may be considered case by case.

Nice to have

  • Experience at a frontier lab, data company, third-party benchmark organization, or public benchmark project.
  • Open-source contributions.

Culture & Benefits

  • High-ownership environment with fast execution and direct responsibility for shipped work.
  • Competitive equity and generous paid time off, including a three-month paid sabbatical after four years.
  • Paid parental leave, wellness stipend, learning and development budget, and team offsites.
  • Medical, dental, vision, life, disability, mental-health, fertility, and family-building benefits.
  • Retirement and pre-tax benefits, including 401(k), HSA, FSA, commuter benefits, and Canadian Group RRSP options.
  • Office-specific perks in San Francisco and Toronto, including meals, transportation support, and transit benefits.

Hiring process

  • Application review focused on a benchmark, leaderboard, evaluation, or dataset previously built.
  • Introductory, technical, workflow, and founder conversations.
  • Paid approximately 40-hour work trial followed by a fast decision.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →