Назад
Company hidden
2 дня назад

SRE L1 Support/Cloud Platform Ops Engineer (AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
SRE L1 Support/Cloud Platform Ops Engineer (AI): Monitoring and supporting GPU cloud infrastructure across compute clusters, networks, storage systems, and environmental sensors with an accent on incident response, hardware triage, and structured escalation. Focus on executing remediation runbooks during the 8AM–8PM PST shift, collecting diagnostics for L2 teams, and converting novel incidents into AIOps automations.

Location: Remote within San Jose, CA or Austin, TX; 8AM–8PM PST shift with rotation and structured handoffs to the APAC operations team

Company

hirify.global provides Bitcoin mining infrastructure and AI computational and cloud infrastructure, operating data centers across multiple countries.

What you will do

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors through centralized dashboards.
  • Respond to alerts and execute runbooks for GPU errors, link flaps, node failures, and storage incidents.
  • Perform hardware triage and standard remediation, including GPU resets, node drains or reboots, link reseating, and BMC recovery.
  • Collect logs, DCGM output, network diagnostics, and hardware health reports for L2 or SME escalation.
  • Manage ServiceNow or Jira incident tickets through resolution or escalation and perform structured shift handoffs.
  • Complete data center tasks such as cabling, hardware replacement, rack and stack, firmware updates, and inventory management.

Requirements

  • 2+ years of experience in NOC, data center operations, or IT support.
  • Basic Linux system administration, including command-line work, log analysis, and service management.
  • Familiarity with monitoring tools such as Prometheus, Grafana, Nagios, or equivalent systems.
  • Experience with ServiceNow or Jira Service Management.
  • Ability to perform physical data center tasks, including rack and stack, cabling, and hardware replacement.
  • Strong communication skills for shift handoffs, incident documentation, and escalation, with curiosity about automation and structured operational data.

Culture & Benefits

  • Work with an AI-operated GPU cloud where incident decisions provide training data for future automations.
  • Maintain operational runbooks so recurring procedures become more executable by the platform.
  • Collaborate with APAC operations through structured 8AM and 8PM PST handoffs.
  • Career growth toward SME roles or platform engineering and automation authorship.
  • Full-time employment with equal employment opportunities in accordance with applicable laws.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →