Назад
2 дня назад

Staff+ Site Reliability Engineer (Safeguards ML Infra)

320 000 - 485 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff+ Site Reliability Engineer (Safeguards ML Infra) (AI/Cloud Infrastructure): Building and operating production infrastructure for Claude’s safety systems with an accent on model-launch safeguards, classifier deployments, and cross-platform configuration validation. Focus on designing repeatable deployment pipelines, automating launch runbooks, detecting configuration drift, and managing incident response for safety-critical systems.

Location: Remote-friendly, with travel required and at least 25% office attendance. Offices are in San Francisco, Seattle, and New York City, United States.

Annual salary: $320,000–$485,000 USD

Company

Anthropic builds reliable, interpretable, and steerable AI systems designed to be safe and beneficial.

What you will do

  • Lead model-release launches by configuring and verifying safeguards and acting as the safeguards point of contact during release windows.
  • Deploy new safety classifiers from research, including canary rollouts, post-deployment validation, and discrepancy investigation.
  • Verify safeguards across first-party infrastructure, AWS Bedrock, GCP Vertex, and other deployment platforms while eliminating configuration drift.
  • Convert launch runbooks and manual checks into tooling, continuous validation, and repeatable self-service deployment pipelines.
  • Build and maintain a safeguards registry recording production deployments, models, platforms, provenance, and deployment ownership.
  • Participate in on-call rotations covering service incidents, model provisioning, and time-sensitive research and safety launches.

Requirements

  • Experience owning production change management at scale, including deployment pipelines, configuration management, and canary analysis.
  • Experience leading high-stakes releases as a launch captain, incident commander, or release owner.
  • Meaningful production on-call and incident-response experience, including postmortem-driven improvements.
  • Hands-on experience deploying and operating AWS and GCP infrastructure at scale.
  • Proficiency in Python.
  • Bachelor’s degree or equivalent education, training, or relevant professional experience.

Nice to have

  • 8+ years of industry software engineering or site reliability engineering experience.
  • Experience reducing operational toil and moving teams from manual deployment processes to self-service pipelines.
  • Experience running launch or production-readiness reviews across multiple teams.
  • Familiarity with LLM inference systems and transformer-based model operations.
  • Rust experience.

Culture & Benefits

  • Collaborative work across research, engineering, policy, and business teams.
  • Flexible working hours and a collaborative office environment.
  • Competitive compensation, equity donation matching, generous vacation, and parental leave.
  • Visa sponsorship is available, subject to role and candidate eligibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →