Назад
Company hidden
6 часов назад

Senior SRE (GPU Infrastructure)

168 000 - 252 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior SRE (GPU Infrastructure) (Linux/AWS/OCI): Operating and scaling production GPU clusters for AI training and inference across on-premises and multi-cloud environments with an accent on Linux performance, high-performance networking, and infrastructure automation. Focus on redesigning clusters for scale, debugging GPU and kernel-level failures, building self-healing systems, and strengthening security and compliance.

Location: Hybrid in Redwood City, California, with remote work available in the US

Salary: $168K–$252K base salary, plus equity

Company

hirify.global builds unified general intelligence that can generate, understand, and operate in the physical world, with a focus on multimodal AI and vision.

What you will do

  • Own production GPU clusters for AI training and inference across on-premises infrastructure, AWS, and OCI.
  • Maintain high availability and performance across thousands of NVIDIA and AMD GPUs.
  • Redesign infrastructure for greater efficiency, reliability, and scale.
  • Tune Linux systems at the OS and kernel level and build automation in Python, Go, or Bash.
  • Act as the final escalation point for GPU, networking, and system failures, collaborating with vendors such as NVIDIA.
  • Strengthen infrastructure security and support SOC 2 Type 1, SOC 2 Type 2, and ISO certifications.

Requirements

  • 5+ years of experience as an SRE, production engineer, or infrastructure engineer in large-scale environments.
  • Deep hands-on expertise with Linux, containerized systems, and low-level performance debugging.
  • Working experience with Terraform, Airflow, and Ray.
  • Strong experience with AWS or OCI.
  • Practical experience with high-performance networking, including InfiniBand, RDMA, or RoCE.
  • Working knowledge of infrastructure security practices and compliance frameworks such as SOC 2 and ISO.

Nice to have

  • Expertise with NVIDIA and AMD GPU tooling, including DCGM and ROCm.
  • Experience managing large-scale GPU clusters for AI/ML training or inference.
  • Familiarity with Kubernetes or orchestration frameworks such as Ray.
  • Expertise in data pipelines and infrastructure.

Culture & Benefits

  • Hands-on, close-to-the-metal work in a fast-paced and less-structured environment.
  • Opportunity to solve complex low-level GPU, networking, Linux, and kernel problems.
  • Base salary of $168K–$252K with equity.
  • Work across on-premises and multi-cloud infrastructure at significant scale.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →