1 день назад
Senior Engineering Manager, Capacity Engineering
405 000 - 485 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Engineering Manager, Capacity Engineering (AI infrastructure): Building and operating data platforms, capacity planning systems, and benchmarking infrastructure for a large-scale AI compute fleet with an accent on utilization, reliability, and efficient resource allocation. Focus on leading senior engineers, setting technical direction, managing production systems, and coordinating capacity and efficiency decisions across infrastructure, research engineering, inference, and finance.
Location: San Francisco, CA; New York City, NY; or Seattle, WA. Hybrid attendance is required at least 25% of the time.
Annual salary: $405,000–$485,000 USD
Company
Anthropic develops reliable, interpretable, and steerable AI systems designed to be safe and beneficial.
What you will do
- Lead, hire, coach, and develop senior and staff-level engineers responsible for capacity engineering.
- Set the roadmap across data platforms, capacity planning, operational assurance, and infrastructure efficiency.
- Oversee pipelines that process occupancy, utilization, billing, and usage telemetry across Kubernetes clusters and cloud providers.
- Build systems for cluster health, capacity planning, allocation monitoring, benchmarking, and workload efficiency.
- Set production standards for Python and SQL systems, data quality, SLOs, incident response, and sustainable on-call practices.
- Partner with infrastructure, inference, research engineering, finance, and senior leadership on capacity decisions, efficiency targets, and spend.
Requirements
- Experience managing software or infrastructure engineering teams, including hiring, performance management, and people development.
- Strong hands-on background in production systems, data engineering, infrastructure, distributed systems, or observability.
- Experience with at least one major cloud provider, Kubernetes-based infrastructure, and observability tools such as Prometheus or Grafana.
- Experience setting and executing engineering roadmaps in ambiguous, high-autonomy environments with multiple stakeholders.
- Excellent communication skills and experience owning on-call and incident management for business-critical systems.
- Bachelor’s degree or equivalent education, training, or professional experience in a relevant field.
Nice to have
- Experience with capacity planning, resource management, product engineering, or FinOps in hyperscale or large-scale ML environments.
- Familiarity with accelerator infrastructure, including GPU metrics, TPU utilization, and ML training or inference systems.
- Experience with multi-cloud billing, telemetry normalization, internal data products, scheduling, packing efficiency, or profiling-driven optimization.
Culture & Benefits
- Collaborative environment focused on large-scale AI research and trustworthy AI systems.
- Flexible working hours and regular research discussions.
- Competitive compensation, optional equity donation matching, generous vacation, and parental leave.
- Visa sponsorship is available, with immigration-lawyer support, although sponsorship is evaluated for each role and candidate.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Manager, Engineering (Data Platform)
310 000 - 385 000$
8 дней назад
Engineering Manager (AI)
210 000 - 300 000$
4 дня назад
Senior Director, Forward Deployed Engineering, Americas (AI)
250 000 - 350 000$
4 дня назад
Engineering Manager, Rep Experience (AI)
259 200 - 324 000$
Lambda
1 день назад
Director of Network Capacity Automation (AI Cloud)
399 000 - 531 000$
5 дней назад
Senior Manager, Data Engineering (AI)
192 100 - 307 800$