Назад
Company hidden
5 дней назад

SRE (AI Inference Infrastructure)

Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
SRE (AI Inference Infrastructure) (Kubernetes/Python/Go): Operating and scaling production infrastructure for a high-performance AI inference service with an accent on releases, capacity changes, cluster upgrades, and observability. Focus on building self-service continuous delivery pipelines, reusable automation, and reliability practices for frontier-class AI workloads.

Location: San Francisco Bay Area or Toronto, United States/Canada

Company

hirify.global builds high-performance AI hardware and infrastructure for training and inference workloads.

What you will do

  • Execute production releases, capacity changes, and cluster upgrades while the SRE function scales.
  • Develop self-service continuous delivery pipelines using Kubernetes, Bazel, Prometheus, Grafana, InfluxDB, Python, and Go.
  • Build reusable automation and internal developer tools to reduce operational toil and cross-team friction.
  • Develop telemetry, observability, and alerting solutions for reliable operations at scale.
  • Collaborate with Cluster Ops and development teams to identify and implement high-impact automation.
  • Contribute to SLOs, post-mortems, reliability metrics, and capacity planning.

Requirements

  • 2–4+ years of SRE experience with a strong operations or automation focus.
  • Production experience with Kubernetes.
  • Proficiency in Python or Go for tools and automation.
  • Experience with Prometheus, Grafana, and observability-driven workflows.
  • Ability to measure and communicate reliability, operational toil, and velocity impact.

Nice to have

  • Hands-on GitOps experience with Argo CD, Flux, or an equivalent tool.
  • Experience building continuous delivery pipelines.
  • Experience with Bazel or similar build systems.
  • Familiarity with capacity planning and on-premises or multi-datacenter environments.

Culture & Benefits

  • Work on a high-performance AI platform and wafer-scale computing architecture.
  • Opportunity to work with cutting-edge AI infrastructure and leading model builders.
  • Direct mentorship from experienced engineers and ownership of production systems.
  • No requirement for 24/7 on-call rotations.
  • Startup vitality, job stability, and a simple, non-corporate work culture.
  • Continuous learning, growth, and support in an inclusive workplace.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →