Назад
9 дней назад

Site Reliability Engineer (AI/GPU)

130 000 - 200 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI/GPU) (AI infrastructure): Building automation, observability, and reliability tooling for production AI and GPU workloads with an accent on Linux, networking, distributed systems, and incident response. Focus on defining SLOs and SLIs, troubleshooting high-scale services, and improving availability, scalability, and efficiency through code.

Location: Houston, New York, San Francisco, or Seattle, United States

Salary: $130,000–$200,000 USD per year, with possible bonus and equity

Company

Nscale provides a GPU cloud infrastructure platform for AI-native startups and global enterprises, from bare metal through platform services.

What you will do

  • Build and own automation and tooling that keeps the platform running and reduces operational toil.
  • Define and maintain SLOs, SLIs, dashboards, metrics, logs, and alerting for service health.
  • Lead incident response, troubleshoot production issues, perform root cause analysis, and run post-incident reviews.
  • Investigate and resolve performance and reliability problems across Linux, networking, and distributed services.
  • Partner with Engineering, Networking, and Infrastructure teams to improve reliability across the stack.
  • Improve availability, scalability, and efficiency through code and participate in the on-call rotation.

Requirements

  • 3–6 years of experience in SRE, systems engineering, or software engineering, including production experience in a data center or cloud environment.
  • Strong programming skills in Python, Go, or a similar language, with a focus on automation.
  • Solid knowledge of Linux, networking fundamentals, and distributed systems.
  • Experience troubleshooting live production issues and owning fixes through post-incident reviews.
  • Fluency with monitoring and observability, including metrics, logs, dashboards, and alerting.
  • Ability to work in a fast-moving environment and participate in an on-call rotation.

Nice to have

  • Experience with AI or GPU workloads or high-performance computing.
  • Familiarity with InfiniBand or RDMA.
  • Experience with Kubernetes and virtualized or bare-metal environments.

Culture & Benefits

  • Ownership, accountability, and direct involvement with production infrastructure.
  • Competitive base salary with equity, reviewed every 12 months.
  • Early scope and a progression plan focused on developing selected skills.
  • Flexible work approach with autonomy over the working day.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →