Назад
обновлено 3 дня назад

Site Reliability Engineer (AI & ML Infrastructure)

136 000 - 240 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI & ML Infrastructure): Build and operate a hybrid infrastructure platform spanning AWS and bare metal data centers for AI/ML research and product development with an accent on Kubernetes, Terraform, and GPU workload orchestration. Focus on architecting scalable, automated, and high-performance environments integrating Slurm and managing complex hybrid cloud infrastructure.

Location: USA - Remote

Base salary: $136,000–$240,000 annually, depending on work location, skills, and experience. Equity and bonuses are offered separately.

Company

Deepgram provides real-time voice AI APIs and self-hosted software for speech-to-text, text-to-speech, and production voice agents.

What you will do

  • Architect and operate Kubernetes-based computing platforms across AWS and on-premises bare-metal data centers.
  • Build reproducible, versioned infrastructure with Terraform and other Infrastructure-as-Code practices.
  • Design AI/ML workload scheduling and orchestration, integrating Slurm with Kubernetes for GPU-intensive workloads.
  • Manage high-performance GPU servers, hybrid networking, service mesh, and storage solutions including CNI, CSI, and S3.
  • Develop monitoring, logging, tracing, operational automation, incident response, and performance-tuning systems.
  • Collaborate with AI researchers and ML engineers to build self-service tools and automate the lifecycle of single-tenant managed deployments.

Requirements

  • Must be based in the USA.
  • 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering.
  • Hands-on experience building and managing production infrastructure with Terraform.
  • Expert knowledge of Kubernetes architecture and operations in large-scale environments.
  • Strong scripting and automation skills with Python, Go, or Bash.
  • Experience with CI/CD systems such as GitLab CI, Jenkins, or ArgoCD, and with developer tooling.

Nice to have

  • Experience with Slurm and high-performance computing for GPU-intensive AI workloads.
  • Experience managing bare-metal infrastructure, including PXE boot, MAAS, provisioning, and lifecycle management.
  • Knowledge of FinOps and cloud cost optimization.
  • Experience with Kubernetes networking and storage technologies such as Calico, Cilium, Ceph, or Rook.
  • Experience in multi-region or hybrid cloud environments.

Culture & Benefits

  • AI use and experimentation are core expectations for every team member.
  • Work focuses on rapidly evolving AI technologies and continuous learning.
  • Infrastructure is treated as a product, with emphasis on developer and researcher experience.
  • Equity and performance bonuses are offered in addition to base salary.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →