Назад
Company hidden
3 дня назад

SRE L1 Support/Cloud Platform Ops Engineers

Формат работы
onsite
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
SRE L1 Support/Cloud Platform Ops Engineers (GPU Cloud/Data Center Operations): Monitoring GPU clusters, networks, storage, and environmental sensors while responding to infrastructure incidents and performing hardware remediation with an accent on Linux operations, diagnostics, and physical data center support. Focus on executing runbooks, collecting escalation data, managing incident tickets, and maintaining reliable 12-hour shift handoffs.

Location: Singapore, SG; on-site data center tasks required

Company

hirify.global provides Bitcoin mining solutions, ASIC hardware, data center infrastructure, and AI cloud capabilities.

What you will do

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors through centralized dashboards.
  • Respond to alerts and execute runbooks for GPU errors, link flaps, node failures, and storage incidents.
  • Perform hardware triage and standard remediation, including GPU resets, node drains and reboots, cable reseating, and BMC recovery.
  • Collect logs, DCGM output, network diagnostics, and hardware health reports for L2 or SME escalation.
  • Manage incident tickets from creation through resolution or escalation in ServiceNow or Jira.
  • Perform rack-and-stack, cabling, labeling, hardware swaps, firmware updates, and inventory tasks.

Requirements

  • 2+ years of experience in NOC, data center operations, or IT support.
  • Basic Linux system administration, including command-line work, log analysis, and service management.
  • Familiarity with Prometheus, Grafana, Nagios, or equivalent monitoring tools.
  • Experience with ServiceNow or Jira Service Management.
  • Ability to perform physical data center work, including rack and stack, cabling, and hardware replacement.
  • Ability to work 8AM–8PM PST shifts with a 12-hour rotation and structured 8AM and 8PM handoffs.

Culture & Benefits

  • Inclusive environment that values authenticity and diverse perspectives.
  • Opportunity to contribute to digital asset and AI cloud infrastructure projects.
  • Autonomy, personal accountability, and opportunities for growth and learning.
  • Training, mentoring, welfare benefits, and developmental opportunities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →