3 дня назад
Senior Platform Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Platform Reliability Engineer (AI): Operating GPU compute fleets, exabyte-scale storage, shared platform services, and observability infrastructure with an accent on production reliability, performance analysis, and guarded automation. Focus on diagnosing internals-level faults, building self-healing remediation, leading major incident recovery, and maintaining service levels across large-scale AI infrastructure.
Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
Company
Technologies develops and operates energy-efficient AI infrastructure, including the AI Cloud and AI FactoryOS platforms, across Asia Pacific.
What you will do
- Operate the multi-tenant control and management plane, including tenancy, quota, and access controls.
- Run exabyte-scale distributed filesystems, high-performance storage, and S3-compatible object storage to defined service levels.
- Operate the production GPU fleet, including node health, fault handling, firmware and driver baselines, remediation, and hardware replacement workflows.
- Build guarded automation, remediation tooling, provisioning frameworks, and CI/CD processes that support self-healing operations.
- Diagnose and tune performance across Linux, networking, software-defined storage, RDMA, GPU Direct Storage, RoCE, and InfiniBand data paths.
- Lead major-incident recovery, production-readiness reviews, patching and vulnerability remediation, vendor escalations, runbook development, and senior-engineer mentoring.
Requirements
- 8+ years of infrastructure, systems, or platform engineering experience, including ownership of production storage and shared infrastructure services in a 24/7 environment.
- Extensive experience with scale-out or parallel storage such as VAST, WEKA, Ceph, Lustre, GPFS, or NetApp.
- Expert-level Linux knowledge covering storage and filesystem internals, kernel and driver behavior, networking, memory, I/O, and performance analysis.
- Experience operating virtualization platforms, observability infrastructure, Kubernetes, infrastructure-as-code, GitOps, and operational automation using Python, Go, or Bash.
- Experience with major-incident response, on-call escalation, post-incident reviews, vendor escalation, runbooks, least privilege, secrets management, certificates, and audited privileged operations.
- Must be based in Australia or Singapore and able to travel to Australian AI Factory sites as required.
Nice to have
- GPU or HPC infrastructure operations experience, including distributed training and inference environments.
- Experience with multi-tenant cloud, service-provider, colocation, GPU-enabled Kubernetes, or bare-metal provisioning environments.
- Knowledge of hardware lifecycle management, DPUs, SmartNICs, InfiniBand, RoCE, policy-as-code, software supply-chain controls, ISO 27001, or SOC 2.
- Bachelor's degree in computer science, engineering, or a related discipline, or equivalent experience and training.
Culture & Benefits
- Permanent full-time employment.
- Participation in a shared after-hours escalation roster within a 24/7 operations function.
- High degree of autonomy, broad technical direction, and direct access to decision makers.
- Opportunity to work on large-scale, sustainable AI infrastructure and GPU systems.
- Inclusive workplace welcoming candidates from diverse backgrounds.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Platform Support Engineers (AI)
4 дня назад
DevOps/Cloud/Platform Engineer
10 дней назад
SRE Monitoring Platform Software Engineer (Early Career / Temporary) (AI Infrastructure)
Airwallex
8 дней назад
Staff Site Reliability Engineer (Fintech)
4 дня назад
DevOps/Cloud/Platform Engineer (Medical Technology)
Canva
7 дней назад