Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Reliability Engineer (HPC) (AI/HPC): Building and operating highly reliable distributed HPC infrastructure for AI model training and inference with an accent on availability, observability, automation, and secure operations. Focus on operating GPU clusters, leading incident response, improving fault tolerance, and optimizing large-scale infrastructure costs and capacity.
Location: Mountain View, United States. Employees living within 50 miles of a designated U.S. Microsoft office are expected to work from that office at least four days per week.
Salary: USD $119,800–$234,700 per year for IC4 roles, or USD $142,800–$274,800 per year for IC5 roles. San Francisco Bay Area and New York City ranges may differ.
Company
Microsoft AI builds artificial intelligence systems, models, services, and infrastructure designed to amplify human potential.
What you will do
- Ensure uptime, resiliency, and fault tolerance for HPC clusters supporting AI model training and inference.
- Design and maintain monitoring, alerting, and logging for GPU, compute, storage, and networking systems.
- Build automation for deployments, incident response, scaling, and failover across CPU and GPU environments.
- Lead on-call rotations, troubleshoot production incidents, conduct blameless postmortems, and drive reliability improvements.
- Maintain secure, privacy-conscious, and compliant model training and serving environments.
- Partner with ML engineers and platform teams to improve developer workflows and research-to-production delivery.
Requirements
- Bachelor’s degree in Computer Science or a related technical field, plus 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, or equivalent experience.
- Experience with Kubernetes, Docker, container orchestration, and CI/CD pipelines for ML training or inference workloads.
- Experience with public cloud platforms such as Azure, AWS, or GCP and infrastructure as code.
- Programming or scripting experience in Python, Go, or Bash.
- Experience with distributed systems, networking, storage, HPC, GPU clusters, and workload schedulers.
- Experience with ML training or inference pipelines, capacity planning, and GPU infrastructure cost optimization.
Nice to have
- Master’s degree in Computer Science or a related technical discipline.
- Experience with Grafana, Datadog, OpenTelemetry, or similar observability tools.
Culture & Benefits
- Work on AI infrastructure supporting large-scale model training and inference.
- Collaborate with product, ML engineering, and platform teams serving users worldwide.
- Employment may include benefits and additional compensation.
- Inclusive equal opportunity workplace with reasonable accommodation support during the application process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Site Reliability Engineering (SRE) Manager
106 000 - 130 600$
6 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
5 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
3 дня назад
Sr. Site Reliability Engineer (AI)
6 дней назад
Site Reliability Engineer
90 000 - 110 000$
3 дня назад
Principal Site Reliability Engineer (Kubernetes)
190 000 - 220 000$