Назад
Company hidden
2 дня назад

Infrastructure Automation Engineer (AI Cluster Commissioning)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Infrastructure Automation Engineer (AI Cluster Commissioning): Developing and operating automation tools for AI cluster commissioning from initial power-on through testing, benchmarking, and operational handover with an accent on GPU infrastructure, Linux systems, networking, storage, monitoring, and Infrastructure as Code. Focus on integrating hardware and cloud tooling, building diagnostics and CI/CD pipelines, and reducing commissioning time and deployment errors across project sites.

Location: Based in Australia or Singapore, with regular visits to project sites in Australia and Southeast Asia; international and domestic travel for on-site deployments and commissioning is required as needed.

Company

hirify.global Technologies develops and operates energy-efficient AI infrastructure and the hirify.global AI Cloud platform across Asia Pacific.

What you will do

  • Develop and maintain custom tools for hardware bring-up, configuration, firmware updates, system inventory, testing, and issue tracking.
  • Implement monitoring, health-check, diagnostic, and remediation tools for GPU, network, storage, and AI infrastructure.
  • Build automation for network deployment, operating-system installation, firmware updates, network configuration, and storage validation.
  • Integrate messaging, inventory, issue-tracking, reporting, dashboard, and operational analytics platforms.
  • Develop and execute project-specific testing and benchmarking tools for commissioning and customer acceptance.
  • Apply Infrastructure as Code, CI/CD, security practices, documentation, and operational handover processes.

Requirements

  • Bachelor’s degree in computer science, engineering, or a related technical field.
  • 5+ years of experience developing software tools, automation platforms, operational tooling, or infrastructure automation for large-scale Linux, cloud, HPC, or AI environments.
  • Strong knowledge of high-performance GPU infrastructure, high-end networking, high-performance storage, and Linux server, network, and storage configuration.
  • Experience developing production tools and APIs, integrating REST APIs, and building automated testing frameworks.
  • Experience with Bash, Python, Ansible, Infrastructure as Code, Redfish, IPMI, DCGM, Prometheus, Grafana, or NetBox.
  • Familiarity with Slurm, NVIDIA DCGM, NCCL, Kubernetes, distributed-system validation, and secure handling of credentials, keys, and secrets.

Culture & Benefits

  • Work with founders and specialists in AI infrastructure, energy systems, and next-generation compute.
  • Founder-led environment with fast decision-making, accessible leadership, and limited bureaucracy.
  • Early ownership and opportunities to grow into new technical domains.
  • Exposure to large-scale AI infrastructure projects across Australia and Southeast Asia.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →