Назад
Company hidden
17 часов назад

Staff SRE (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff SRE (AI): Building self-service reliability platforms and GitOps-driven delivery systems for high-performance AI inference infrastructure across datacenters and cloud environments with an accent on observability, automation, and SLO-driven operations. Focus on architecting model release pipelines, capacity provisioning, cluster upgrades, and reliability practices for latency, throughput, and accuracy at scale.

Location: Hybrid in the SF Bay Area or Toronto

Company

hirify.global builds wafer-scale AI hardware and high-performance inference infrastructure for model labs, enterprises, and AI-native startups.

What you will do

  • Define and implement strategies for reliable software delivery and operations across multiple datacenters and cloud environments.
  • Architect self-service platforms and internal tooling for product teams, external customers, and cluster operators.
  • Build GitOps-driven delivery for model releases, capacity provisioning, and cluster upgrades.
  • Define SLOs, SLIs, error budgets, blameless postmortems, chaos testing, and capacity forecasting for inference workloads.
  • Mentor SREs, support critical incident escalations, and prioritize automation based on production pain points.
  • Measure toil reduction, deployment velocity, SLO compliance, MTTR, and self-service adoption.

Requirements

  • 8+ years of experience in SRE, infrastructure engineering, or platform engineering.
  • Experience improving automation and reliability at large scale in demanding technical environments.
  • Deep expertise operating large heterogeneous clusters with a proprietary cloud control plane.
  • Experience designing and delivering CI/CD or GitOps systems with Argo CD or similar tools.
  • Hands-on experience with Loki, Tempo, Mimir, Prometheus, or comparable observability systems.
  • Ability to lead complex projects, influence stakeholders, and communicate technical direction clearly.

Nice to have

  • Production experience with Bazel or other large-scale build systems.
  • Experience with AI/ML inference, model serving runtimes, GPU or wafer-scale orchestration, and latency or accuracy SLOs.
  • Experience with predictive autoscaling, chaos engineering, or cost-aware capacity planning.

Culture & Benefits

  • Work on a breakthrough AI platform and high-performance AI supercomputer.
  • Opportunities to publish and open-source AI research.
  • Startup vitality combined with job stability.
  • Simple, non-corporate work culture that respects individual beliefs.
  • No 24/7 on-call rotation is required.
  • Commitment to an inclusive and diverse work environment with continuous learning and support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →