Назад
обновлено 4 дня назад

Product Designer, Evals & Prompts (AI)

305 000 - 385 000$
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Product Designer, Evals & Prompts (AI): Building evaluation systems, prompt tooling, and internal interfaces for Claude with an accent on LLM eval pipelines, automated graders, and model-release workflows. Focus on designing low-code evaluation tools, scaling a harness across 50–100 tools, and distinguishing model regressions from harness failures.

Location: San Francisco, CA; hybrid policy requiring staff to be in an Anthropic office at least 25% of the time

Annual salary: $305,000–$385,000 USD

Company

Anthropic builds reliable, interpretable, and steerable AI systems intended to be safe and beneficial for users and society.

What you will do

  • Write, test, revise, and ship prompts for Claude tools, features, and product behaviors.
  • Build automated graders, comparison sets, regression suites, and evaluation pipelines for LLM products.
  • Create visual, low-code evaluation tools that allow designers to compare prompt variants and model outputs without engineering support.
  • Support model releases by testing product surfaces, writing prompt migrations, and measuring regressions.
  • Build and scale a test harness for 50–100 tools, including sandboxed tool calls and reproducible settings.
  • Convert evaluation findings into training signals such as graders, human-feedback questions, and preference pairs.

Requirements

  • Production-quality Python.
  • Experience building and maintaining LLM evaluation pipelines, including graders, rubrics, comparison sets, and regression suites.
  • Experience creating internal tools with interfaces for non-technical users.
  • Experience building test harnesses, sandboxing tool calls, and pinning settings for comparable runs.
  • Experience shipping prompts or working closely with prompt engineers, with an understanding of model-to-model prompt behavior.
  • Bachelor’s degree or equivalent education, training, or experience in a relevant field.

Nice to have

  • Experience working within a model-launch cycle.
  • A/B testing experience and the ability to connect offline evaluations with online outcomes.
  • Front-end or notebook-to-app experience.
  • Experience converting product rubrics into training signals.

Culture & Benefits

  • Collaborative work across research, engineering, policy, business, and product teams.
  • Visa sponsorship is available, with immigration lawyer support, subject to role and candidate eligibility.
  • Flexible working hours, generous vacation and parental leave, and competitive benefits.
  • Optional equity donation matching and an office designed for collaboration.
  • Frequent research discussions focused on steerable and trustworthy AI.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →