8 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Senior Platform Engineer (AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Senior Platform Engineer (AI): Building MLOps and Kubernetes-based platform infrastructure for a large-scale, energy-efficient AI cloud with an accent on infrastructure automation, observability, security, and NVIDIA GPU networking. Focus on scaling platform engineering capabilities, integrating NVIDIA software services, developing self-service tools, and leading incident response and root-cause analysis.
Location: Singapore
Company
Technologies develops and operates energy-efficient AI infrastructure, including the AI Cloud GPU platform, by combining AI software orchestration, liquid cooling, energy management, and construction.
What you will do
- Build MLOps capabilities for reproducible, scalable, and secure machine-learning workflows across internal and customer-facing environments.
- Improve the DevOps platform, including CI/CD and infrastructure-service integrations, for reliability, scalability, and security.
- Design, implement, operate, and secure Kubernetes production infrastructure supporting NVIDIA GB300 NVL72 systems with Quantum-X800 InfiniBand or Spectrum-X Ethernet.
- Develop and scale observability platforms using telemetry, OpenTelemetry, and related monitoring technologies.
- Integrate central services with NVIDIA Mission Control, NETQ, UFM, and NMX, while enhancing internal self-service platform products.
- Lead incident response, participate in on-call rotation, and conduct root-cause analyses to improve operational maturity.
Requirements
- Bachelorβs degree in computer science or a related technical field.
- 7+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer, MLOps Engineer, or Observability Engineer.
- Strong experience with infrastructure-as-code, configuration management, CI/CD, Terraform, Ansible, GitHub Actions, Jenkins, or ArgoCD.
- Experience with Docker, Kubernetes networking and cluster management, observability stacks, telemetry solutions, OpenTelemetry, and compliance automation.
- Competent scripting and programming skills in Bash, Python, or Go, plus knowledge of Linux internals, networking stacks, and distributed storage.
- Clear and effective English communication, written and spoken.
Nice to have
- Experience in high-growth startups or regulated industries with strong security and data-privacy requirements.
- Experience with SOC 2 Type 2 and ISO 27001.
Culture & Benefits
- Work at the intersection of sustainability, artificial intelligence, and next-generation infrastructure.
- Collaborate closely with founders and experienced engineers in an emerging company.
- Contribute to making AI compute more sustainable, accessible, and affordable.
- Inclusive environment focused on diverse backgrounds and authentic collaboration.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
1 Π΄Π΅Π½Ρ Π½Π°Π·Π°Π΄
Staff Platform Engineer, AI Agent Infrastructure & Security
SandboxAQ
6 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Staff Platform Engineer (AI)
121Β 600 - 228Β 000$
7 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Cloud Platform Engineer
9 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Platform Engineer (AI)
IG
6 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Platform Engineer (DevOps)
16 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄