Назад
3 дня назад

Evals Infrastructure Tech Lead / Manager (AI)

500 000 - 850 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Evals Infrastructure Tech Lead / Manager (AI): Building and scaling distributed systems to measure AI model performance and safety with an accent on orchestration, reproducibility, and high-throughput inference. Focus on designing compute allocation strategies and building sandboxed execution infrastructure for agentic evals.

Location: Hybrid in San Francisco, CA (minimum 25% office attendance required)

Salary: $500,000 - $850,000 USD

Company

Anthropic is a public benefit corporation dedicated to creating reliable, interpretable, and steerable AI systems that are safe and beneficial for society.

What you will do

  • Lead the team building distributed systems that schedule, orchestrate, and execute evals for frontier model training.
  • Optimize eval throughput and cost through efficient compute allocation and accelerator pool queuing.
  • Develop and scale harnesses that researchers use to define, run, and iterate on evaluations.
  • Ensure evaluation reliability through determinism, reproducibility, and rigorous uncertainty quantification.
  • Integrate eval signals into dashboards and reviews to inform critical launch decisions.
  • Combine direct engineering contributions with team management, prioritization, and coaching.

Requirements

  • Proven track record of leading end-to-end technical projects on large-scale distributed systems.
  • 1+ years of experience managing engineers or serving as a tech lead with direct reports.
  • Strong proficiency in Python and Rust.
  • Experience building high-throughput, fault-tolerant systems on cloud or on-prem accelerator fleets.
  • Must be based in or able to work from the San Francisco office under the hybrid policy.

Nice to have

  • Experience with LLM inference or training infrastructure.
  • Knowledge of eval or benchmarking systems, particularly agentic evals requiring sandboxed execution.
  • Statistical literacy regarding variance, confidence intervals, and sample-size sufficiency.
  • Experience with observability and regression detection over time-series metrics.

Culture & Benefits

  • Competitive compensation and optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours and a collaborative office environment.
  • Visa sponsorship support for eligible candidates.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →