Назад
Company hidden
12 часов назад

Site Reliability Engineer (AI Infrastructure)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI Infrastructure): Supporting the deployment, maintenance, and reliability of GPU servers, storage systems, networking equipment, and software for AI-accelerated HPC infrastructure with an accent on hardware diagnostics, Linux troubleshooting, firmware lifecycle management, and secure operations. Focus on deploying bare-metal systems and HPC clusters, troubleshooting Kubernetes and Slurm environments, responding to incidents through a 24/7 on-call rotation, and improving operational documentation.

Location: Launceston, Tasmania, Australia

Company

hirify.global develops and operates energy-efficient AI infrastructure, including liquid-cooled AI Factories and the hirify.global AI Cloud GPU platform, across the Asia-Pacific region.

What you will do

  • Deploy, configure, and maintain high-end GPU servers, storage servers, networking equipment, and software components in secure environments.
  • Perform hardware diagnostics, systems checks, firmware updates, and firmware compliance activities.
  • Support customer environments involving bare-metal systems, HPC clusters, Kubernetes, and Slurm.
  • Provide first-line engineering support for onsite hardware, network, and software incidents.
  • Troubleshoot incidents, escalate critical issues, and provide technical support to the Global Operations Centre.
  • Participate in a 24/7 on-call rotation and maintain accurate incident and operational documentation.

Requirements

  • Bachelor’s degree in computer engineering, computer science, or a related technical field.
  • 5+ years of experience in field service technical roles.
  • Strong knowledge of server hardware, firmware lifecycles, Linux environments, hardware troubleshooting, and physical and system-level security.
  • Experience with Bash or Python scripting.
  • Familiarity with configuration management, CI/CD, workload management, cluster software, and observability tools such as Slurm, Kubernetes, NVIDIA BCM, Prometheus, Grafana, or ELK.
  • Strong problem-solving, analytical, communication, and teamwork skills.

Culture & Benefits

  • Work alongside founders and specialists in AI infrastructure, energy systems, and next-generation computing.
  • Gain exposure to AI infrastructure through an NVIDIA Cloud and Engineering partner in the Asia-Pacific region.
  • Operate in a founder-led environment with fast decision-making, accessible leadership, and limited bureaucracy.
  • Take ownership of work that contributes directly to the stability and growth of large-scale AI infrastructure.
  • Participate in team meetings and knowledge-sharing sessions focused on collaboration and continuous learning.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →