Назад
обновлено 1 день назад

Senior Site Reliability Engineer (AI Infrastructure Operations)

170 000 - 265 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AI Infrastructure Operations) (AI infrastructure): Owning reliability for critical production services and automation across GPU cloud infrastructure with an accent on SLOs, observability, incident response, and scalable systems engineering. Focus on designing resilient distributed systems, leading complex incidents, building toil-reducing automation, and improving Kubernetes, Linux, networking, and bare-metal operations.

Location: Houston, San Francisco, or Seattle

Salary: $170,000–$265,000 USD base salary per year, with potential bonus and equity.

Company

Nscale operates a GPU cloud for AI-native startups and global enterprises, providing infrastructure from bare metal through platform services.

What you will do

  • Own reliability for critical production services and set reliability direction across the platform.
  • Define and maintain SLOs, incident processes, observability, alerting, and sustainable on-call practices.
  • Lead architecture and design reviews to build reliability into systems from the beginning.
  • Lead complex production incidents, identify root causes, and implement lasting fixes.
  • Build tooling and automation that reduces operational toil across the SRE team.
  • Mentor SREs through design reviews, pairing, and incident debriefs while raising engineering standards.

Requirements

  • 6–10 years of experience in SRE, systems engineering, or software engineering with production ownership at scale.
  • Strong software engineering skills in Python, Go, or a similar language.
  • Deep knowledge of Linux, networking, distributed systems, Kubernetes, and virtualized or bare-metal environments.
  • Experience running AI or GPU workloads or high-performance computing environments.
  • Experience implementing SLOs, observability, large-scale alerting, incident management, and sustainable on-call practices.
  • Track record of serving as a senior technical voice during incidents and design reviews and mentoring other engineers.

Nice to have

  • Familiarity with high-performance networking, including InfiniBand and RDMA.

Culture & Benefits

  • Shared on-call rotation with a focus on reducing operational load over time.
  • Direct ownership of reliability practices across the platform.
  • Flexible approach to organizing the workday.
  • Competitive base salary with compensation reviewed every 12 months.
  • Potential bonus, equity, medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →