Назад
1 день назад

Senior Engineering Manager, Capacity Engineering

405 000 - 485 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Engineering Manager, Capacity Engineering (AI infrastructure): Building and operating data platforms, capacity planning systems, and benchmarking infrastructure for a large-scale AI compute fleet with an accent on utilization, reliability, and efficient resource allocation. Focus on leading senior engineers, setting technical direction, managing production systems, and coordinating capacity and efficiency decisions across infrastructure, research engineering, inference, and finance.

Location: San Francisco, CA; New York City, NY; or Seattle, WA. Hybrid attendance is required at least 25% of the time.

Annual salary: $405,000–$485,000 USD

Company

Anthropic develops reliable, interpretable, and steerable AI systems designed to be safe and beneficial.

What you will do

  • Lead, hire, coach, and develop senior and staff-level engineers responsible for capacity engineering.
  • Set the roadmap across data platforms, capacity planning, operational assurance, and infrastructure efficiency.
  • Oversee pipelines that process occupancy, utilization, billing, and usage telemetry across Kubernetes clusters and cloud providers.
  • Build systems for cluster health, capacity planning, allocation monitoring, benchmarking, and workload efficiency.
  • Set production standards for Python and SQL systems, data quality, SLOs, incident response, and sustainable on-call practices.
  • Partner with infrastructure, inference, research engineering, finance, and senior leadership on capacity decisions, efficiency targets, and spend.

Requirements

  • Experience managing software or infrastructure engineering teams, including hiring, performance management, and people development.
  • Strong hands-on background in production systems, data engineering, infrastructure, distributed systems, or observability.
  • Experience with at least one major cloud provider, Kubernetes-based infrastructure, and observability tools such as Prometheus or Grafana.
  • Experience setting and executing engineering roadmaps in ambiguous, high-autonomy environments with multiple stakeholders.
  • Excellent communication skills and experience owning on-call and incident management for business-critical systems.
  • Bachelor’s degree or equivalent education, training, or professional experience in a relevant field.

Nice to have

  • Experience with capacity planning, resource management, product engineering, or FinOps in hyperscale or large-scale ML environments.
  • Familiarity with accelerator infrastructure, including GPU metrics, TPU utilization, and ML training or inference systems.
  • Experience with multi-cloud billing, telemetry normalization, internal data products, scheduling, packing efficiency, or profiling-driven optimization.

Culture & Benefits

  • Collaborative environment focused on large-scale AI research and trustworthy AI systems.
  • Flexible working hours and regular research discussions.
  • Competitive compensation, optional equity donation matching, generous vacation, and parental leave.
  • Visa sponsorship is available, with immigration-lawyer support, although sponsorship is evaluated for each role and candidate.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →