обновлено 8 дней назад
Safeguards Enforcement Analyst (User Well-being)
20 417 - 23 750$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Safeguards Enforcement Analyst (AI Safety): Designing and deploying mental health guardrails for AI systems with an accent on detection systems, intervention evaluation, and policy enforcement. Focus on translating clinical guidance into measurable rubrics and optimizing detection models to prevent harms like self-harm and emotional dependence.
Location: San Francisco, CA; New York City, NY; or Washington, DC. Hybrid work is required, with staff expected to be in an office at least 25% of the time.
Annual salary: $245,000–$285,000 USD
Company
Anthropic develops reliable, interpretable, and steerable AI systems designed to be safe and beneficial for users and society.
What you will do
- Design and execute mental health safeguard interventions, define success metrics, and curate evaluation datasets.
- Partner with Engineering and Data Science to build, tune, and validate detection models, including threshold-setting and precision-recall tradeoffs.
- Monitor intervention and detection-system performance over time and review flagged content to improve enforcement policies.
- Support product features that connect users with crisis resources, collaborating with Product, Legal, and external partners on referral pathways and user-facing content.
- Identify policy gaps from real-world scenarios and incorporate emerging AI policy and mental-health research into workflows.
Requirements
- Experience in trust and safety, product policy, content moderation, or a related field involving mental health, suicide, self-harm, or other well-being harms.
- Experience designing experiments, evaluations, or measurement studies and translating policy definitions into measurable rubrics or classification criteria.
- Experience coordinating content-review operations, quality assurance, and workflow management.
- Proficiency in SQL or other data-analysis tools, plus experience with generative AI products and effective prompts for review, classification, or evaluation.
- Strong judgment in ambiguous, high-consequence cases and ability to communicate emerging risks to cross-functional stakeholders.
- Bachelor’s degree or equivalent education, training, and experience in a relevant field.
Nice to have
- Subject-matter expertise in mental health through academia, clinical practice, crisis intervention, trust and safety, or a related setting.
- Experience building or evaluating LLM-based classification systems.
- Experience using agentic tools such as Claude Code to scale analysis or automate recurring work.
- Experience working in crisis support.
Culture & Benefits
- Collaborative environment focused on large-scale research and trustworthy AI.
- Flexible working hours and an office-based collaboration model.
- Competitive compensation, optional equity donation matching, generous vacation, and parental leave.
- Visa sponsorship may be available, with immigration-lawyer support, depending on the role and candidate.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
14 дней назад
AI Model Assurance Analyst
98 800 - 196 000$
11 дней назад
Review Operations Leader (AI)
160 000 - 190 000$
11 дней назад
GenAI Policy Specialist
96 000 - 217 023$
8 дней назад
Operations Specialist (AI)
55 000 - 75 000$
13 дней назад
Senior Human Data Operations Partner, AI
13 дней назад
SMB AI Power User — Competitive Evaluations (AI)
40 - 50$