3 дня назад
Site Reliability Engineer (AI Accelerator Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI Accelerator Infrastructure): Building and operating reliable infrastructure across colocation facilities, on-premises AI lab clusters, cloud platforms, and customer-facing services with an accent on infrastructure automation, observability, and high-speed interconnects. Focus on designing Terraform and Ansible workflows, troubleshooting bare-metal and Kubernetes environments, and resolving P0/P1 incidents through root-cause analysis and permanent reliability improvements.
Location: Santa Clara, United States; hybrid
Company
designs and manufactures purpose-built AI inference silicon and develops the infrastructure supporting silicon engineering, AI/ML research, and customer deployments.
What you will do
- Own the reliability and availability of colocation server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services.
- Provision and troubleshoot bare-metal servers, operating systems, networks, storage, and physical hardware.
- Operate high-speed interconnect environments including InfiniBand, RoCE, and high-speed Ethernet.
- Build infrastructure-as-code and automation with Terraform and Ansible for provisioning, configuration management, fleet health checks, auto-remediation, and self-service tooling.
- Design monitoring dashboards, alerting, and service-level indicators with Prometheus, Grafana, or DataDog; participate in on-call incident response.
- Produce root-cause analyses, improve system reliability, and maintain customer-environment documentation and operational runbooks.
Requirements
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience, plus 5+ years in SRE, infrastructure engineering, or systems administration.
- Strong Linux systems knowledge covering networking, storage, systemd, package management, kernel parameters, and performance diagnostics.
- Hands-on experience with colocation or on-premises server infrastructure, physical hardware, rack networking, and bare-metal provisioning.
- Production experience writing and maintaining Terraform and/or Ansible configurations, along with Kubernetes operations covering troubleshooting, workloads, storage, and networking.
- Experience with Prometheus and Grafana or DataDog, Python and/or Bash automation, structured incident response, and RCA production.
- Ability to own infrastructure systems, document implementations, and operate effectively in a fast-moving startup environment.
Nice to have
- Customer-facing infrastructure, cloud operations across AWS, Azure, or GCP, and hybrid cloud/on-premises environments.
- HPC scheduler experience with Slurm, LSF, or an equivalent platform.
- Knowledge of InfiniBand, RoCE, or NVLink configuration and troubleshooting.
- Go programming for SRE tooling and experience with large-scale automation, fleet auto-healing, or AIOps-driven operations.
Culture & Benefits
- Six-month contract with potential conversion to a full-time position.
- Hands-on, high-ownership infrastructure role spanning colocation, on-premises labs, and cloud platforms.
- Collaborative and inclusive environment emphasizing respect, humility, direct communication, and execution.
- Equal opportunity workplace committed to a welcoming and empowered work environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Site Reliability Engineer - Vice President
Lambda
1 день назад
Senior Site Reliability Engineer (Kubernetes)
267 000 - 356 000$
3 дня назад
Production Engineer (AI Infrastructure)
172 000 - 209 000$
Nscale
1 день назад
Infrastructure Software Engineer (AI)
150 000 - 215 000$
3 дня назад
Software Engineer (AI Infrastructure)
3 дня назад