2 дня назад
Infrastructure Automation Engineer (AI Cluster Commissioning)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Automation Engineer (AI Cluster Commissioning): Developing and operating automation tools for AI cluster commissioning from initial power-on through testing, benchmarking, and operational handover with an accent on GPU infrastructure, Linux systems, networking, storage, monitoring, and Infrastructure as Code. Focus on integrating hardware and cloud tooling, building diagnostics and CI/CD pipelines, and reducing commissioning time and deployment errors across project sites.
Location: Based in Australia or Singapore, with regular visits to project sites in Australia and Southeast Asia; international and domestic travel for on-site deployments and commissioning is required as needed.
Company
Technologies develops and operates energy-efficient AI infrastructure and the AI Cloud platform across Asia Pacific.
What you will do
- Develop and maintain custom tools for hardware bring-up, configuration, firmware updates, system inventory, testing, and issue tracking.
- Implement monitoring, health-check, diagnostic, and remediation tools for GPU, network, storage, and AI infrastructure.
- Build automation for network deployment, operating-system installation, firmware updates, network configuration, and storage validation.
- Integrate messaging, inventory, issue-tracking, reporting, dashboard, and operational analytics platforms.
- Develop and execute project-specific testing and benchmarking tools for commissioning and customer acceptance.
- Apply Infrastructure as Code, CI/CD, security practices, documentation, and operational handover processes.
Requirements
- Bachelor’s degree in computer science, engineering, or a related technical field.
- 5+ years of experience developing software tools, automation platforms, operational tooling, or infrastructure automation for large-scale Linux, cloud, HPC, or AI environments.
- Strong knowledge of high-performance GPU infrastructure, high-end networking, high-performance storage, and Linux server, network, and storage configuration.
- Experience developing production tools and APIs, integrating REST APIs, and building automated testing frameworks.
- Experience with Bash, Python, Ansible, Infrastructure as Code, Redfish, IPMI, DCGM, Prometheus, Grafana, or NetBox.
- Familiarity with Slurm, NVIDIA DCGM, NCCL, Kubernetes, distributed-system validation, and secure handling of credentials, keys, and secrets.
Culture & Benefits
- Work with founders and specialists in AI infrastructure, energy systems, and next-generation compute.
- Founder-led environment with fast decision-making, accessible leadership, and limited bureaucracy.
- Early ownership and opportunities to grow into new technical domains.
- Exposure to large-scale AI infrastructure projects across Australia and Southeast Asia.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Cluster Administration Engineer (AI)
200 000 - 400 000SGD
6 дней назад
Senior Software DevOps Engineer (DevOps)
4 дня назад
DevOps Engineer (Blockchain)
4 дня назад
DevOps Engineer (AWS/Kubernetes)
110 000 - 140 000AUD
4 дня назад
Lead DevOps Engineer (Azure/Kubernetes)
4 дня назад