Назад
2 месяца назад

Red Team Engineer (AI Safeguards)

320 000 - 405 000$
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Red Team Engineer (AI Safeguards): Conduct adversarial testing across deployed AI systems and product surfaces to uncover vulnerabilities and emergent abuse before malicious actors exploit them with an accent on full kill-chain attack simulation, novel testing for agent/tool use capabilities, and building automated testing frameworks at scale. Focus on chaining multiple exploitation techniques, researching non-obvious prompt injection vectors, and translating findings into concrete safety improvements with measurable detection effectiveness.

Location: Remote-Friendly (Travel Required) | San Francisco, CA

Salary: $320,000–$405,000 USD (annual)

Company

Anthropic builds reliable, interpretable, and steerable AI systems with a focus on safety and beneficial outcomes.

What you will do

  • Run comprehensive adversarial testing across product surfaces, creating creative multi-technique attack scenarios.
  • Research and implement novel testing approaches for emerging AI capabilities (agent systems, tool use, and new interaction paradigms).
  • Design and execute full kill-chain attacks that emulate real-world threat actors pursuing specific malicious objectives.
  • Build and maintain systematic testing methodologies covering every aspect of systems, including continuous assessment at scale via automated testing frameworks.
  • Collaborate with Product, Engineering, and Policy teams to turn findings into concrete improvements and define metrics for detection effectiveness.

Requirements

  • Experience in penetration testing, red teaming, or application security.
  • Experience with model jailbreaking and testing large-scale agentic workflows for non-obvious prompt injection vectors.
  • Strong web application security skills with hands-on security testing tools (e.g., Burp Suite, Metasploit, custom scripting frameworks).
  • Experience building custom automation, including LLM-specific testing frameworks.
  • Proven track record of discovering novel attack vectors and chaining vulnerabilities creatively (e.g., CVEs, blog posts, or disclosed bug bounty reports).
  • Strong written and verbal communication skills to explain technical concepts to varied audiences.

Nice to have

  • Experience with AI/ML security or adversarial machine learning.
  • Understanding of AI safety beyond traditional security, including modern guardrails against jailbreaks.
  • Experience testing API security and rate-limiting systems.
  • Background in anti-fraud, trust & safety, or abuse prevention systems; authorization bypass and business logic vulnerabilities.
  • Familiarity with distributed systems, infrastructure security, and abuse detection mechanisms (including engineering novel bypasses).

Culture & Benefits

  • Location-based hybrid policy: expect to be in one of the offices at least 25% of the time (some roles may require more).
  • Visa sponsorship available; reasonable efforts made to support offers where possible.
  • Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, and flexible working hours.
  • Collaborative office environment and frequent research discussions.

Hiring process

  • Application review with emphasis on safety-relevant technical communication.
  • Interviews and evaluations focused on red teaming, adversarial testing, and ability to translate findings into improvements.
  • Use of AI in the application process follows the company’s candidate AI guidance policy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →