Назад
Company hidden
7 часов назад

AI Red Team Engineer

100 000 - 200 000GBP
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Red Team Engineer (AI monitoring and coding-agent security): Building automated red-teaming pipelines and campaigns to identify attack surfaces in AI monitors, including Watcher and frontier labs' monitoring systems, with an accent on adversarial testing, monitor evasion, and agent failure modes. Focus on designing scalable attack campaigns, escalating red-team/blue-team games, and translating findings into actionable fixes, publications, and monitoring improvements.

Location: In-person role based in the London or San Francisco office, with work-from-home arrangements available

Salary: £100,000–£200,000 per year, approximately $150,000–$270,000 USD

Company

hirify.global develops AI safety research, monitoring systems, and the Watcher coding-agent security product, focusing on risks from loss of control and deceptive AI behavior.

What you will do

  • Design and run red-teaming campaigns using failure-mode injection, static monitoring benchmarks, and dynamic off-policy control testing.
  • Identify novel attack surfaces and monitor-evasion techniques that have not yet been tested.
  • Build and maintain automated pipelines for attacking AI monitors at scale.
  • Design iterative red-team/blue-team exercises with control researchers to increase attack difficulty as monitors improve.
  • Translate campaign findings into actionable recommendations for monitor developers and contribute to external publications and internal reports.
  • Feed research findings into Watcher development and Apollo's blue-teaming work.

Requirements

  • At least 2 years of experience in offensive security, adversarial machine learning, or AI red-teaming.
  • Strong experience using, comparing, or developing frontier AI coding agents.
  • Experience designing and executing structured adversarial testing campaigns.
  • Strong Python programming skills.
  • Strong written communication skills for clear and credible publications and campaign reports.
  • Ability to work independently on open-ended adversarial problems.

Nice to have

  • Familiarity with AI safety concepts, especially agent-related risks.
  • Experience with LLM-as-a-judge systems or AI monitoring.
  • Background in penetration testing, CTFs, or broader computer security.

Culture & Benefits

  • Flexible working hours and schedule, with work-from-home arrangements.
  • Unlimited vacation and sick leave.
  • Up to six months of paid parental leave.
  • Comprehensive health, dental, and vision insurance, plus retirement savings with employer matching.
  • Provided meals and snacks on workdays, paid work trips, and a $1,000 annual professional development budget.
  • Relocation support and visa fees where applicable.

Hiring process

  • Screening interview, followed by a three-hour take-home test.
  • Three technical interviews and a final interview with the CEO.
  • No LeetCode-style general coding interviews; AI tools may be used for the take-home test.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →