Назад
Company hidden
3 дня назад

Site Reliability Engineer (AI Accelerator Infrastructure)

Формат работы
hybrid
Тип работы
project
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI Accelerator Infrastructure): Building and operating reliable infrastructure across colocation facilities, on-premises AI lab clusters, cloud platforms, and customer-facing services with an accent on infrastructure automation, observability, and high-speed interconnects. Focus on designing Terraform and Ansible workflows, troubleshooting bare-metal and Kubernetes environments, and resolving P0/P1 incidents through root-cause analysis and permanent reliability improvements.

Location: Santa Clara, United States; hybrid

Company

hirify.global designs and manufactures purpose-built AI inference silicon and develops the infrastructure supporting silicon engineering, AI/ML research, and customer deployments.

What you will do

  • Own the reliability and availability of colocation server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services.
  • Provision and troubleshoot bare-metal servers, operating systems, networks, storage, and physical hardware.
  • Operate high-speed interconnect environments including InfiniBand, RoCE, and high-speed Ethernet.
  • Build infrastructure-as-code and automation with Terraform and Ansible for provisioning, configuration management, fleet health checks, auto-remediation, and self-service tooling.
  • Design monitoring dashboards, alerting, and service-level indicators with Prometheus, Grafana, or DataDog; participate in on-call incident response.
  • Produce root-cause analyses, improve system reliability, and maintain customer-environment documentation and operational runbooks.

Requirements

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience, plus 5+ years in SRE, infrastructure engineering, or systems administration.
  • Strong Linux systems knowledge covering networking, storage, systemd, package management, kernel parameters, and performance diagnostics.
  • Hands-on experience with colocation or on-premises server infrastructure, physical hardware, rack networking, and bare-metal provisioning.
  • Production experience writing and maintaining Terraform and/or Ansible configurations, along with Kubernetes operations covering troubleshooting, workloads, storage, and networking.
  • Experience with Prometheus and Grafana or DataDog, Python and/or Bash automation, structured incident response, and RCA production.
  • Ability to own infrastructure systems, document implementations, and operate effectively in a fast-moving startup environment.

Nice to have

  • Customer-facing infrastructure, cloud operations across AWS, Azure, or GCP, and hybrid cloud/on-premises environments.
  • HPC scheduler experience with Slurm, LSF, or an equivalent platform.
  • Knowledge of InfiniBand, RoCE, or NVLink configuration and troubleshooting.
  • Go programming for SRE tooling and experience with large-scale automation, fleet auto-healing, or AIOps-driven operations.

Culture & Benefits

  • Six-month contract with potential conversion to a full-time position.
  • Hands-on, high-ownership infrastructure role spanning colocation, on-premises labs, and cloud platforms.
  • Collaborative and inclusive environment emphasizing respect, humility, direct communication, and execution.
  • Equal opportunity workplace committed to a welcoming and empowered work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →