10 часов назад
Senior AI Infrastructure Engineer (Virtualisation)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Infrastructure Engineer (Virtualisation) (AI infrastructure, storage, and Kubernetes): Designing, building, and operating software-defined infrastructure for large-scale AI workloads with an accent on virtualisation, bare-metal provisioning, high-performance storage, and GPU clusters. Focus on developing multi-tenant control planes, validating RDMA-enabled storage performance, and building Kubernetes operators and orchestration frameworks for reliable AI infrastructure.
Location: Singapore or Australia, including Melbourne, Sydney, or Launceston
Company
develops software-defined infrastructure and sustainable solutions for large-scale AI workloads.
What you will do
- Design and implement scalable, multi-tenant control planes for AI and infrastructure workloads.
- Develop and operate exabyte-scale S3-compatible object storage, distributed file systems, and high-performance filesystems.
- Provision and manage bare-metal infrastructure using platforms such as Base Command Manager, Warewulf, Ironic, and MaaS.
- Work with RDMA, GPU Direct Storage, RoCE, InfiniBand, DPDK, Ceph, Weka, DAOS, Kubernetes, and composable storage clusters.
- Monitor, debug, benchmark, and optimise internal clusters and storage platforms in collaboration with SRE, operations, and networking teams.
- Build automation, validation, CI/CD, Kubernetes operators, and orchestration frameworks for large-scale GPU cluster commissioning.
Requirements
- 6–10 years of experience in infrastructure engineering and/or storage engineering.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- Hands-on experience with bare-metal provisioning and software-defined storage platforms such as Ceph, Weka, Vast Data, DAOS, or Lustre.
- Strong knowledge of Linux systems engineering, Kubernetes, cloud-native infrastructure, distributed systems, networking, and high-performance environments.
- Experience with automation tools such as Ansible, Helm, Terraform/OpenTofu, or equivalent, and programming in Go, Bash, Rust, or Python.
- Experience supporting production services through an on-call rotation and documenting architecture, procedures, and performance results.
Culture & Benefits
- Full-time employment.
- Collaboration across engineering, operations, SRE, site operations, and networking teams.
- Focus on continuous technical improvement, knowledge transfer, and innovation in AI and HPC infrastructure.
- Commitment to diversity and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Infrastructure Operations Engineer (APAC) (AI)
165 000 - 205 000SGD
12 часов назад
Infrastructure Engineer (AI)
10 часов назад
Cloud Platform Engineer
10 часов назад
Staff AI Scheduling & Orchestration Engineer (AI)
10 часов назад
Senior Staff DevOps Engineer – Orchestration (AI)
10 часов назад