Назад
2 дня назад

Safeguards Enforcement Analyst (User Well-being)

245 000 - 285 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Safeguards Enforcement Analyst (AI Safety): Designing and deploying mental health guardrails for AI systems with an accent on detection systems, intervention evaluation, and policy enforcement. Focus on translating clinical guidance into measurable rubrics and optimizing detection models to prevent harms like self-harm and emotional dependence.

Location: Must be based in the United States; Remote-Friendly with a hybrid policy requiring office presence in San Francisco, New York City, or Washington, DC at least 25% of the time.

Salary: $245,000 - $285,000 USD

Company

Anthropic is a public benefit corporation dedicated to creating reliable, interpretable, and steerable AI systems that are safe and beneficial for society.

What you will do

  • Design and execute interventions, define key metrics, and curate evaluation datasets for mental health guardrails.
  • Partner with Engineering and Data Science to build, tune, and validate detection models, managing precision and recall tradeoffs.
  • Review flagged content to drive enforcement and identify policy gaps based on real-world scenarios.
  • Develop in-product features to connect users to crisis resources in collaboration with Product and Legal teams.
  • Monitor the performance of detection systems and interventions over time to ensure efficacy.
  • Stay updated on emerging AI policy and mental health research to inform decision-making and workflows.

Requirements

  • Experience in Trust & Safety, product policy, or content moderation, specifically focusing on mental health, suicide, and self-harm harms.
  • Proficiency in SQL and other data analysis tools to measure intervention efficacy and monitor workflow health.
  • Experience designing experiments, evaluations, or measurement studies to determine intervention success.
  • Ability to translate complex policy definitions into measurable rubrics, review guidelines, or classification criteria.
  • Experience with generative AI products, including writing effective prompts for content review and evaluation.
  • Must be based in the United States.

Nice to have

  • Subject matter expertise in mental health from academia, clinical practice, or crisis intervention.
  • Experience building or evaluating LLM-based classification systems.
  • Experience using agentic tools (e.g., Claude Code) to scale analysis or automate recurring tasks.
  • Professional experience working within crisis support environments.

Culture & Benefits

  • Collaborative "big science" environment focusing on high-impact, large-scale research efforts.
  • Competitive compensation with optional equity donation matching.
  • Generous vacation and parental leave policies.
  • Flexible working hours and access to high-quality office spaces.
  • Visa sponsorship availability for eligible candidates.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →