Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Red Team Engineer (AI Safeguards): Conduct adversarial testing across deployed AI systems and product surfaces to uncover vulnerabilities and emergent abuse before malicious actors exploit them with an accent on full kill-chain attack simulation, novel testing for agent/tool use capabilities, and building automated testing frameworks at scale. Focus on chaining multiple exploitation techniques, researching non-obvious prompt injection vectors, and translating findings into concrete safety improvements with measurable detection effectiveness.
Location: Remote-Friendly (Travel Required) | San Francisco, CA
Salary: $320,000–$405,000 USD (annual)
Company
Anthropic builds reliable, interpretable, and steerable AI systems with a focus on safety and beneficial outcomes.
What you will do
- Run comprehensive adversarial testing across product surfaces, creating creative multi-technique attack scenarios.
- Research and implement novel testing approaches for emerging AI capabilities (agent systems, tool use, and new interaction paradigms).
- Design and execute full kill-chain attacks that emulate real-world threat actors pursuing specific malicious objectives.
- Build and maintain systematic testing methodologies covering every aspect of systems, including continuous assessment at scale via automated testing frameworks.
- Collaborate with Product, Engineering, and Policy teams to turn findings into concrete improvements and define metrics for detection effectiveness.
Requirements
- Experience in penetration testing, red teaming, or application security.
- Experience with model jailbreaking and testing large-scale agentic workflows for non-obvious prompt injection vectors.
- Strong web application security skills with hands-on security testing tools (e.g., Burp Suite, Metasploit, custom scripting frameworks).
- Experience building custom automation, including LLM-specific testing frameworks.
- Proven track record of discovering novel attack vectors and chaining vulnerabilities creatively (e.g., CVEs, blog posts, or disclosed bug bounty reports).
- Strong written and verbal communication skills to explain technical concepts to varied audiences.
Nice to have
- Experience with AI/ML security or adversarial machine learning.
- Understanding of AI safety beyond traditional security, including modern guardrails against jailbreaks.
- Experience testing API security and rate-limiting systems.
- Background in anti-fraud, trust & safety, or abuse prevention systems; authorization bypass and business logic vulnerabilities.
- Familiarity with distributed systems, infrastructure security, and abuse detection mechanisms (including engineering novel bypasses).
Culture & Benefits
- Location-based hybrid policy: expect to be in one of the offices at least 25% of the time (some roles may require more).
- Visa sponsorship available; reasonable efforts made to support offers where possible.
- Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, and flexible working hours.
- Collaborative office environment and frequent research discussions.
Hiring process
- Application review with emphasis on safety-relevant technical communication.
- Interviews and evaluations focused on red teaming, adversarial testing, and ability to translate findings into improvements.
- Use of AI in the application process follows the company’s candidate AI guidance policy.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
AI Red Team Engineer
100 000 - 200 000GBP
8 дней назад
Staff Security Researcher (AI Security)
168 000 - 238 000$
12 дней назад
Sr. AI Red Team Engineer
145 900 - 234 200$
CrowdStrike
10 дней назад
Red Team Services Consultant (Cybersecurity)
95 000 - 140 000$
13 дней назад
Red Team Manager (AI Security)
170 000 - 230 000$
11 дней назад
Insider Threat Counter Operations (ITCO) Engineer (Cybersecurity)
142 000 - 309 000$