Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Cyber Evaluations Engineer (AI): Building and running evaluations that measure cyber-relevant model capabilities, safeguard robustness, and production abuse detection with an accent on offensive security, adversarial data analysis, and evaluation tooling. Focus on designing cyber misuse probes, translating policy lines into layered detection architecture, and measuring detection precision and coverage over time.
Location: San Francisco, CA or Washington, DC; hybrid attendance at one of the offices at least 25% of the time
Annual salary: $300,000–$405,000 USD
Company
Anthropic develops reliable, interpretable, and steerable AI systems designed to be safe and beneficial for users and society.
What you will do
- Design and run capability, uplift, and safety evaluations for cyber-relevant risks in new models.
- Execute safeguard-robustness testing before major model releases.
- Analyze evaluation results and communicate findings to technical, policy, and cross-functional stakeholders.
- Design, prototype, and tune probes for detecting cyber misuse in production.
- Work with cyber policy and engineering partners to build layered abuse-detection architecture and improve safeguards.
- Build and maintain internal tooling for running and scoring evaluations.
Requirements
- Experience building or running evaluations, benchmarks, or test suites for software or ML systems on short, fixed timelines.
- Hands-on cybersecurity experience, such as CTF participation, vulnerability research, exploit development, or security research.
- Proficiency in Python.
- Strong communication skills for presenting evaluation results to cross-functional and policy stakeholders.
- Bachelor’s degree or equivalent education, training, or experience in a relevant field.
Nice to have
- Deep offensive-security or security-research experience, including AI security benchmark development.
- Experience analyzing adversarial or abuse data, AI/ML evaluation frameworks, or pre-release software and model testing.
- Experience authoring detection content such as Sigma, YARA, Suricata, or SIEM rules, or building ML-based abuse detection.
- Experience with coordinated vulnerability disclosure practices or onsite government testing engagements.
- Active secret security clearance or eligibility to obtain one.
Culture & Benefits
- Collaborative research environment focused on large-scale efforts in trustworthy and steerable AI.
- Flexible working hours.
- Generous vacation and parental leave.
- Competitive compensation, benefits, and optional equity donation matching.
- Visa sponsorship is available, subject to role and candidate eligibility.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →