Назад
1 день назад

Staff + Sr. Software Engineer, Cloud Inference Launch Engineering (AI)

320 000 - 485 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering (AI): Building validation and CI/CD infrastructure for inference servers and load balancers across AWS, GCP, Azure, and future cloud platforms with an accent on correctness, performance, reliability, and cost-efficient accelerator usage. Focus on launching frontier models, integrating inference features, diagnosing cross-cloud discrepancies, and designing shadow-traffic, performance-baseline, and regression-checking systems.

Location: San Francisco, CA or Seattle, WA. Hybrid work is expected, with staff working from one of the offices at least 25% of the time; some roles may require more office attendance.

Annual salary: $320,000–$485,000 USD

Company

Anthropic is a public benefit corporation building reliable, interpretable, and steerable AI systems.

What you will do

  • Bring up inference for new model architectures and launch frontier models on cloud platforms in sync with the first-party platform.
  • Integrate inference features such as structured sampling and prompt caching across AWS, GCP, Azure, and future cloud service providers.
  • Identify and fix cross-platform inference gaps involving configuration drift, observability, deployment patterns, and provider-specific behavior.
  • Design and own CI/CD infrastructure for inference servers and load balancers, including shadow traffic, throughput and latency baselines, and correctness checks.
  • Reduce merge-to-production cycle time through faster, more parallel, and cost-effective validation on constrained accelerator capacity.
  • Analyze observability data to identify performance bottlenecks, cost anomalies, and production regressions.

Requirements

  • Significant software engineering experience with high-performance, large-scale distributed systems serving millions of users.
  • Experience building automation or test infrastructure that improved release velocity or reliability.
  • Experience operating services on AWS, GCP, or Azure, with exposure to Kubernetes, infrastructure as code, or container orchestration.
  • Strong interest in LLM serving; prior inference or machine learning experience is not required.
  • Ability to collaborate across internal teams and external partners, learn new technologies quickly, and own problems end to end.
  • Bachelor’s degree or equivalent education, training, or relevant experience.

Nice to have

  • LLM inference optimization, batching, and caching experience.
  • Experience with capacity-constrained scheduling or shared-resource test infrastructure.
  • Understanding of multi-region deployments, request routing, load balancing, and global traffic management.
  • Experience scaling infrastructure with cloud service provider partner teams across networking, security, privacy, and managed services.
  • Proficiency in Python or Rust.

Culture & Benefits

  • Collaborative work across research, engineering, policy, and business functions.
  • Flexible working hours and an office environment for collaboration.
  • Generous vacation and parental leave.
  • Competitive compensation and optional equity donation matching.
  • Visa sponsorship is available, with reasonable efforts and immigration-lawyer support for eligible offers.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →