Назад
Company hidden
1 день назад

Site Reliability Engineer (AI Infrastructure)

175 000 - 265 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI Infrastructure): Building and operating reliable infrastructure across colocation facilities, on-premises GPU clusters, cloud environments, and customer-facing platform services with an accent on automation, observability, and bare-metal operations. Focus on designing infrastructure as code, resolving incidents across the hardware-to-application stack, and supporting AI and HPC workloads with Kubernetes, cloud platforms, and high-performance interconnects.

Location: Santa Clara, United States; hybrid work

Salary: $175,000–$265,000 per year, plus equity and bonus opportunities

Company

hirify.global develops software and hardware infrastructure for generative AI and focuses on collaborative, execution-oriented engineering.

What you will do

  • Own reliability and availability across colocation server fleets, on-premises GPU clusters, cloud environments, and customer-facing platform services.
  • Provision and operate infrastructure from bare metal through Kubernetes, including operating systems, networking, storage, hardware, and autoscaling.
  • Automate provisioning, deployment, operations, host lifecycle management, fleet health checks, remediation, and networking using Terraform and/or Ansible.
  • Design monitoring, alerting, dashboards, SLIs, and AIOps-driven detection workflows with tools such as Prometheus, Grafana, DataDog, and Splunk.
  • Participate in on-call rotation, resolve incidents from bare metal to application layers, and produce root-cause analyses for P0/P1 incidents.
  • Support CI/CD, QA, HPC workloads, operational runbooks, capacity planning, hardware lifecycle management, and cloud cost optimization.

Requirements

  • 7+ years of experience in SRE, infrastructure engineering, or systems administration.
  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • Strong Linux expertise with colocation or on-premises infrastructure, including networking, storage, systemd, kernel parameters, performance diagnostics, rack networking, and bare-metal provisioning.
  • Production experience writing and maintaining Terraform and/or Ansible configurations.
  • Operational Kubernetes experience covering cluster troubleshooting, workloads, storage, and networking.
  • Experience with observability tools, production-quality Python and/or Bash scripting, incident response, structured triage, and RCA follow-through.

Nice to have

  • Experience with customer-facing infrastructure and external reliability commitments.
  • Cloud operations across AWS, Azure, or GCP in hybrid cloud and on-premises environments.
  • Production experience with AIOps, intelligent alerting, anomaly detection, or LLM-assisted diagnostics.
  • Experience with Slurm, LSF, InfiniBand, RoCE, NVLink, or large-scale infrastructure automation.

Culture & Benefits

  • Collaborative and inclusive environment built around respect, humility, direct communication, and diverse perspectives.
  • Medical, dental, and vision coverage.
  • 401(k) and an employee rewards program focused on personal and family wellbeing.
  • Equity and performance-based bonus opportunities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →