Назад
Company hidden
5 дней назад

Principal SRE (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal SRE (AI): Architecting self-service delivery, observability, capacity orchestration, rollout safety, and operational automation for large-scale AI inference infrastructure with an accent on reliability, fleet management, and production control planes. Focus on building unified capacity management systems, designing SLO-based reliability practices, and enabling safe self-service workflows across datacenters and cloud environments.

Location: On-site in the SF Bay Area or Toronto

Company

hirify.global Systems develops wafer-scale AI hardware and high-performance inference infrastructure for model labs, enterprises, and AI-native startups.

What you will do

  • Define and implement strategies for reliably delivering and operating software across multiple datacenters and cloud environments.
  • Architect self-service platforms and internal tooling for safely triggering and observing critical workflows.
  • Design reliability practices for inference workloads, including SLOs, SLIs, error budgets, postmortems, chaos testing, and capacity forecasting.
  • Support incident escalations, mentor senior SREs, and prioritize automation based on production pain points.
  • Measure impact through toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.

Requirements

  • 15+ years of experience in SRE, infrastructure engineering, or platform engineering in large-scale production environments.
  • Deep experience with compute fleets, control planes, schedulers, orchestration systems, capacity management, and reliability automation.
  • Experience driving cross-team architecture for production control planes, fleet management, capacity orchestration, or self-service infrastructure platforms.
  • Ability to lead complex technical programs, influence senior stakeholders, mentor engineers, and communicate technical strategy.
  • Hands-on experience with observability, incident response, and SLO-based reliability management across metrics, logs, traces, alerting, and dashboards.

Nice to have

  • Experience with Bazel or other large-scale build systems.
  • Background in AI/ML inference systems, model serving, disaggregated inference, GPU orchestration, latency and accuracy SLOs, or drift monitoring.
  • Experience with predictive autoscaling, chaos engineering, or cost-aware capacity management.

Culture & Benefits

  • Work on a wafer-scale AI platform designed beyond traditional GPU constraints.
  • Opportunities to publish and open-source AI research.
  • Access to one of the fastest AI supercomputers in the world.
  • Startup vitality combined with job stability.
  • Non-corporate culture focused on individual beliefs, learning, growth, and support.
  • The role does not require 24/7 on-call rotations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →