Назад
5 часов назад

Engineering Manager, Inference Infrastructure (AI)

405 000 - 625 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Engineering Manager, Inference Infrastructure (AI): Leading the control plane for Anthropic's inference fleet and request path with an accent on load balancing, capacity coordination, distributed systems, and operational reliability. Focus on designing fleet-control strategies across heterogeneous hardware and multiple cloud providers, measuring throughput, latency, utilization, and cost improvements, and building reliable on-call and incident-response practices.

Location: San Francisco, CA; New York City, NY; or Seattle, WA. Hybrid policy: staff must work from one of the offices at least 25% of the time.

Annual salary: $405,000–$625,000 USD

Company

Anthropic builds reliable, interpretable, and steerable AI systems designed to be safe and beneficial.

What you will do

  • Own the technical roadmap for coordinating the inference fleet, including traffic routing, capacity placement, cache placement, demand response, and control-plane protocols.
  • Lead ML platform, infrastructure, and distributed-systems engineers responsible for the inference request path.
  • Partner with product, inference engine, performance, and capacity teams to deliver measurable throughput, latency, utilization, and cost improvements.
  • Set technical strategy across heterogeneous hardware, multiple cloud providers, and all serving surfaces.
  • Run on-call rotations, incident response, postmortems, and deployment safety practices.
  • Develop and retain engineers, hire to a high technical bar, coach teams, and shape team structure as the scope grows.

Requirements

  • Engineering management experience leading critical-path production infrastructure teams at scale.
  • Deep systems experience in areas such as load balancing, scheduling, cluster orchestration, autoscaling, distributed state, or high-performance networking.
  • Experience shipping and quantifying performance or efficiency improvements, including cost impact.
  • Experience operating production infrastructure with on-call, incident response, capacity events, and deployment discipline.
  • Ability to build strong relationships across teams and make architectural decisions in complex systems.
  • Bachelor’s degree or equivalent education, training, or professional experience; curiosity about machine learning systems and transformer inference.

Nice to have

  • 5+ years of engineering management experience.
  • Experience with LLM inference serving, including KV caching, continuous batching, request scheduling, or prefill/decode disaggregation.
  • Experience with Kubernetes internals, cluster schedulers, autoscalers, load balancers, service meshes, or fleet control planes.
  • Experience operating across multiple clouds, partner platforms, heterogeneous accelerator fleets, supercomputing, or hyperscaler-scale infrastructure.
  • Experience leading multiple teams through rapid growth, hiring, onboarding, and team restructuring.

Culture & Benefits

  • Collaborative environment focused on large-scale AI research and long-term impact.
  • Competitive compensation and benefits, including optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours and office collaboration.
  • Visa sponsorship is available, with immigration-lawyer support, though sponsorship cannot be guaranteed for every role or candidate.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →