Назад
Company hidden
4 дня назад

Head of Site Reliability Engineering (AI)

195 000 - 285 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
head
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Head of Site Reliability Engineering (AI) (AI infrastructure): Building and leading the SRE function for AI inference silicon infrastructure across colocation facilities, on-premises labs, cloud environments, and customer-facing platform services with an accent on reliability engineering, observability, infrastructure automation, and storage architecture. Focus on defining SLOs, designing incident response and on-call systems, operating hybrid-cloud infrastructure, and leading the migration to enterprise-grade shared storage.

Location: Santa Clara, United States; hybrid

Salary: $195,000–$285,000 per year, plus equity and bonus opportunities

Company

hirify.global develops purpose-built AI inference silicon and the software and infrastructure that support generative AI applications.

What you will do

  • Build and lead the SRE function, define its charter, establish SRE practices, and hire and develop a team of 1–3 SRE engineers.
  • Direct Data Center & Lab Technician operations across on-premises and colocation facilities.
  • Own 24×7 reliability for colocation, on-premises lab clusters, cloud environments, and customer-facing platform services.
  • Define SLIs, SLOs, error budgets, on-call rotations, incident management, and root-cause analysis processes.
  • Own observability, infrastructure-as-code automation, self-healing systems, FinOps, capacity planning, and workload placement.
  • Lead the migration from ad-hoc JBOD storage to an enterprise-grade shared storage platform across on-premises, colocation, and cloud environments.

Requirements

  • 15+ years of experience in SRE, infrastructure engineering, or production engineering, plus 5+ years leading SRE or infrastructure engineering teams.
  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • Deep Linux expertise, including bare-metal operations, NAS/SAN, NFS/SMB, snapshots, replication, and hybrid-cloud storage.
  • Production-scale Terraform and Ansible experience, Kubernetes operations, and hands-on operation of colocation and on-premises hardware.
  • Experience with observability platforms such as Prometheus, Grafana, Datadog, or Splunk, plus strong Python and/or Go scripting skills.
  • Ability to communicate infrastructure health and operational risk to senior and non-technical stakeholders and build structure in ambiguous environments.

Nice to have

  • Experience with customer-facing infrastructure, multi-cloud hybrid operations across AWS, Azure, and GCP, or FinOps.
  • Knowledge of InfiniBand, RoCE, NVLink, HPC schedulers such as Slurm or LSF, and structured incident or change management frameworks.
  • Technical writing, conference presentations, or open-source contributions in reliability, observability, or HPC infrastructure.

Culture & Benefits

  • Collaborative and inclusive environment focused on respect, humility, direct communication, and execution.
  • Medical, dental, and vision coverage.
  • 401(k) and an employee wellbeing-focused rewards package.
  • Equity and performance-based bonus opportunities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →