Назад
обновлено 2 месяца назад

Staff + Senior Software Engineer, Cloud Inference Launch Engineering (AI)

320 000 - 485 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff + Senior Software Engineer (Cloud Inference/AI): Building and optimizing the validation pipeline and CI/CD infrastructure for LLM inference servers across major cloud providers with an accent on scalability, correctness, and release velocity. Focus on designing high-performance distributed systems, reducing merge-to-production cycle time, and solving cross-platform infrastructure discrepancies.

Location: Hybrid in San Francisco, CA (at least 25% office presence required)

Salary: $320,000 - $485,000 USD

Company

Anthropic is a public benefit corporation dedicated to creating reliable, interpretable, and steerable AI systems.

What you will do

  • Lead frontier model launches by bringing up inference for new architectures across cloud platforms.
  • Integrate new inference features like structured sampling and prompt caching into production.
  • Design and own CI/CD infrastructure for inference servers and load balancers with shadow traffic and performance baselines.
  • Optimize merge-to-production cycle time by improving validation speed and cost-effectiveness.
  • Analyze observability data to identify and remediate performance bottlenecks and cost anomalies.

Requirements

  • Significant software engineering experience with high-performance, large-scale distributed systems.
  • Proven track record of building automation or test infrastructure to improve release velocity.
  • Experience with at least one major cloud platform (AWS, GCP, or Azure), Kubernetes, and Infrastructure as Code.
  • Must be based in or able to work hybrid in San Francisco, CA.
  • Bachelor’s degree in a relevant field or equivalent professional experience.

Nice to have

  • Experience with LLM inference optimization, batching, and caching.
  • Knowledge of multi-region deployments, request routing, and global traffic management.
  • Proficiency in Python or Rust.

Culture & Benefits

  • Collaborative "big science" approach to AI research.
  • Competitive compensation and optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours and a collaborative office space.
  • Visa sponsorship availability for qualified candidates.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →