12 часов назад
Site Reliability Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI Infrastructure): Supporting the deployment, maintenance, and reliability of GPU servers, storage systems, networking equipment, and software for AI-accelerated HPC infrastructure with an accent on hardware diagnostics, Linux troubleshooting, firmware lifecycle management, and secure operations. Focus on deploying bare-metal systems and HPC clusters, troubleshooting Kubernetes and Slurm environments, responding to incidents through a 24/7 on-call rotation, and improving operational documentation.
Location: Launceston, Tasmania, Australia
Company
develops and operates energy-efficient AI infrastructure, including liquid-cooled AI Factories and the AI Cloud GPU platform, across the Asia-Pacific region.
What you will do
- Deploy, configure, and maintain high-end GPU servers, storage servers, networking equipment, and software components in secure environments.
- Perform hardware diagnostics, systems checks, firmware updates, and firmware compliance activities.
- Support customer environments involving bare-metal systems, HPC clusters, Kubernetes, and Slurm.
- Provide first-line engineering support for onsite hardware, network, and software incidents.
- Troubleshoot incidents, escalate critical issues, and provide technical support to the Global Operations Centre.
- Participate in a 24/7 on-call rotation and maintain accurate incident and operational documentation.
Requirements
- Bachelor’s degree in computer engineering, computer science, or a related technical field.
- 5+ years of experience in field service technical roles.
- Strong knowledge of server hardware, firmware lifecycles, Linux environments, hardware troubleshooting, and physical and system-level security.
- Experience with Bash or Python scripting.
- Familiarity with configuration management, CI/CD, workload management, cluster software, and observability tools such as Slurm, Kubernetes, NVIDIA BCM, Prometheus, Grafana, or ELK.
- Strong problem-solving, analytical, communication, and teamwork skills.
Culture & Benefits
- Work alongside founders and specialists in AI infrastructure, energy systems, and next-generation computing.
- Gain exposure to AI infrastructure through an NVIDIA Cloud and Engineering partner in the Asia-Pacific region.
- Operate in a founder-led environment with fast decision-making, accessible leadership, and limited bureaucracy.
- Take ownership of work that contributes directly to the stability and growth of large-scale AI infrastructure.
- Participate in team meetings and knowledge-sharing sessions focused on collaboration and continuous learning.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 часов назад
Senior Staff DevOps Engineer – Orchestration (AI)
10 часов назад
Senior Infrastructure Engineer (AI Cloud)
7 дней назад
Customer Reliability Engineer (AWS)
12 часов назад
Platform Engineer (AI)
3 дня назад
Infrastructure Engineer (AI)
208 000 - 269 000$
15 часов назад