Staff+ Site Reliability Engineer (Safeguards ML Infra)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Staff+ Site Reliability Engineer (Safeguards ML Infra): Building and operating production infrastructure for Claude's safety systems with an accent on deployment safety, change management, and continuous validation. Focus on automating model launch workflows, verifying safeguards across AWS, GCP, and first-party platforms, and leading incident response for safety-critical systems.
Location: Remote-friendly, but staff must work from one of the San Francisco, Seattle, or New York City offices at least 25% of the time.
Salary: $405,000–$485,000 USD annual compensation.
Company
Anthropic develops reliable, interpretable, and steerable AI systems designed to be safe and beneficial.
What you will do
- Act as launch captain for model releases by configuring and verifying safeguards across supported platforms.
- Deploy new safety classifiers through canary rollouts, post-deployment validation, and rollback decisions.
- Detect and eliminate configuration drift across first-party infrastructure, AWS Bedrock, and GCP Vertex.
- Convert launch runbooks, manual checks, and one-off deployments into automated tooling and repeatable pipelines.
- Build and maintain a provenance-aware safeguards registry covering models, platforms, deployments, and ownership.
- Participate in on-call rotations for incidents, model provisioning, and time-sensitive research and safety launches.
Requirements
- Deep experience with production change management, deployment pipelines, configuration management, or canary analysis at scale.
- Experience owning high-stakes releases as a launch captain, incident commander, or release owner.
- Strong production on-call and incident response experience, including postmortem-driven improvements.
- Hands-on experience deploying and operating cloud platforms at scale, especially AWS and GCP.
- Proficiency in Python.
- Bachelor's degree or an equivalent combination of education, training, and experience.
Nice to have
- 8+ years of software engineering or site reliability engineering experience.
- Experience reducing operational toil and moving teams from manual deployments to self-service pipelines.
- Experience running launch or production-readiness reviews across multiple teams.
- Familiarity with LLM inference systems and transformer-based model operations.
- Rust experience.
Culture & Benefits
- Collaborative work across research, engineering, policy, and business functions.
- Flexible working hours and a collaborative office environment.
- Generous vacation and parental leave.
- Competitive compensation and benefits with optional equity donation matching.
- Visa sponsorship is available, subject to role and candidate eligibility.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →