Назад
Company hidden
3 часа назад

Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract (AI)

195 000 - 285 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
director
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract (AI): Building and leading the SRE function for colocation facilities, on-premises lab clusters, multi-cloud environments, and customer-facing AI accelerator infrastructure with an accent on reliability architecture, observability, automation, and capacity planning. Focus on establishing SLOs and incident management, designing shared storage and hybrid-cloud platforms, and scaling infrastructure operations across HPC workloads and silicon programs.

Location: Santa Clara, United States; hybrid

Salary: $195,000–$285,000 per year, plus equity and bonus opportunities

Company

hirify.global designs and manufactures purpose-built AI inference silicon and develops the supporting software, hardware, QA, research, and infrastructure systems.

What you will do

  • Build and lead the SRE function, including its charter, technical roadmap, team structure, and operational standards.
  • Hire, develop, and retain a team of 3–5 SRE engineers while directing data center and lab technician operations.
  • Establish SLOs, error budgets, on-call rotations, incident management, observability, and RCA practices from a zero baseline.
  • Own reliability across colocation, on-premises lab clusters, AWS, Azure, GCP, and customer-facing platform services.
  • Drive infrastructure-as-code, self-healing automation, capacity planning, FinOps, and migration to enterprise shared storage.
  • Partner with DevOps, hardware, software, and executive stakeholders on infrastructure reliability and HPC workload requirements.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field, plus 15+ years of SRE, infrastructure, or production engineering experience.
  • 5+ years leading SRE or infrastructure engineering teams, including building or significantly rebuilding an SRE function.
  • Deep Linux, TCP/IP, RDMA, bare-metal, enterprise storage, colocation, on-premises hardware, and Kubernetes experience.
  • Production-scale Terraform and Ansible expertise, including module design, remote state, environment isolation, and change governance.
  • Experience owning Prometheus, Grafana, and/or Datadog observability and writing production services or automation in Python and/or Go.
  • Strong executive communication and the ability to create structure in a high-ambiguity, low-process environment.

Nice to have

  • Experience with customer-facing infrastructure, InfiniBand, RoCE, NVLink, or HPC schedulers such as Slurm and LSF.
  • Multi-cloud hybrid operations, FinOps, ITIL, technical writing, conference talks, or open-source contributions.

Culture & Benefits

  • Collaborative, inclusive culture centered on respect, humility, direct communication, and execution.
  • Medical, dental, and vision coverage.
  • 401(k) and an inclusive rewards plan supporting employee wellbeing and dependents.
  • Six-month contract with the possibility of conversion to a full-time role.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →