обновлено 3 дня назад
Site Reliability Engineer (AI & ML Infrastructure)
136 000 - 240 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI & ML Infrastructure): Build and operate a hybrid infrastructure platform spanning AWS and bare metal data centers for AI/ML research and product development with an accent on Kubernetes, Terraform, and GPU workload orchestration. Focus on architecting scalable, automated, and high-performance environments integrating Slurm and managing complex hybrid cloud infrastructure.
Location: USA - Remote
Base salary: $136,000–$240,000 annually, depending on work location, skills, and experience. Equity and bonuses are offered separately.
Company
Deepgram provides real-time voice AI APIs and self-hosted software for speech-to-text, text-to-speech, and production voice agents.
What you will do
- Architect and operate Kubernetes-based computing platforms across AWS and on-premises bare-metal data centers.
- Build reproducible, versioned infrastructure with Terraform and other Infrastructure-as-Code practices.
- Design AI/ML workload scheduling and orchestration, integrating Slurm with Kubernetes for GPU-intensive workloads.
- Manage high-performance GPU servers, hybrid networking, service mesh, and storage solutions including CNI, CSI, and S3.
- Develop monitoring, logging, tracing, operational automation, incident response, and performance-tuning systems.
- Collaborate with AI researchers and ML engineers to build self-service tools and automate the lifecycle of single-tenant managed deployments.
Requirements
- Must be based in the USA.
- 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering.
- Hands-on experience building and managing production infrastructure with Terraform.
- Expert knowledge of Kubernetes architecture and operations in large-scale environments.
- Strong scripting and automation skills with Python, Go, or Bash.
- Experience with CI/CD systems such as GitLab CI, Jenkins, or ArgoCD, and with developer tooling.
Nice to have
- Experience with Slurm and high-performance computing for GPU-intensive AI workloads.
- Experience managing bare-metal infrastructure, including PXE boot, MAAS, provisioning, and lifecycle management.
- Knowledge of FinOps and cloud cost optimization.
- Experience with Kubernetes networking and storage technologies such as Calico, Cilium, Ceph, or Rook.
- Experience in multi-region or hybrid cloud environments.
Culture & Benefits
- AI use and experimentation are core expectations for every team member.
- Work focuses on rapidly evolving AI technologies and continuous learning.
- Infrastructure is treated as a product, with emphasis on developer and researcher experience.
- Equity and performance bonuses are offered in addition to base salary.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Staff Software Engineer, Cloud Infrastructure (AI)
181 000 - 265 000$
7 дней назад
DevOps Engineer (AI)
90 000 - 130 000$
SandboxAQ
4 дня назад
Staff Platform Engineer (AI)
121 600 - 228 000$
Writer
8 дней назад
Infrastructure Engineer (AI)
173 000 - 240 000$
Airwallex
9 дней назад
Senior Software Engineer (Infrastructure AI)
140 000 - 240 000$
7 дней назад
DevOps Engineer (AWS)
100 000 - 145 000$