19 часов назад
Staff Production Engineer (AI)
209 000 - 253 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Production Engineer (AI): Improving the reliability, scalability, and performance of an energy-efficient GPU cloud platform with an accent on SLOs, incident response, observability, and distributed systems. Focus on designing self-healing automation, strengthening disaster recovery, and solving complex reliability challenges across AI and HPC infrastructure.
Location: Sunnyvale, California, United States; on-site
Salary: $209,000–$253,000 annually plus bonus and restricted stock units.
Company
builds vertically integrated, energy-efficient AI infrastructure and operates an AI-optimized cloud platform for demanding compute workloads.
What you will do
- Define, measure, and improve availability metrics, service level indicators, and service level objectives for the cloud platform.
- Lead production incident response, service disruption resolution, post-incident reviews, and root cause analysis.
- Architect and improve observability using Prometheus, Grafana, Alertmanager, and OpenTelemetry.
- Identify reliability risks and performance bottlenecks across distributed systems and GPU infrastructure.
- Design automation, remediation tooling, and self-healing infrastructure to reduce operational toil and improve recovery times.
- Partner with compute, networking, storage, and platform teams while mentoring engineers and promoting reliability practices.
Requirements
- 8+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations.
- Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems.
- Strong Linux/Unix debugging skills across kernel and user space.
- Knowledge of Kubernetes, distributed systems, virtualization, and cloud platforms such as AWS or GCP.
- Experience with incident management, reliability frameworks, monitoring and observability, and infrastructure-as-code tools such as Terraform or Ansible.
- Proficiency in Go, Python, C, or C++, with strong communication and cross-functional collaboration skills.
Nice to have
- Experience leading Kubernetes or container orchestration platforms at scale.
- Experience with operational readiness reviews, change management, structured root cause analysis, or automated remediation.
- Interest in scaling AI or HPC infrastructure and mentoring Production Engineering teams.
Culture & Benefits
- Health, dental, and vision insurance options, including HDHP and PPO plans.
- Employer HSA contributions, paid parental leave, life insurance, and disability coverage.
- 401(k) with a 100% employer match up to 4% of salary.
- Paid time off, holidays, tuition reimbursement, and a $300 monthly commuter benefit.
- Restricted stock units, cell phone reimbursement, Teladoc, Calm, and MetLife Legal benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
22 часа назад
Forward Deployed Engineers (AI Infrastructure)
180 000 - 240 000$
17 часов назад
Staff SRE (AI)
6 дней назад
Staff Platform Engineer (Kubernetes)
250 000 - 285 000$
22 часа назад
Sr. DevOps Engineer (AI)
175 000 - 195 000$
2 дня назад
Staff Software Engineer (Platform/DevOps)
160 000 - 208 000$
Windsurf
7 дней назад