Назад
Company hidden
8 часов Π½Π°Π·Π°Π΄

Senior Platform Engineer (AI)

Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
Singapore
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Senior Platform Engineer (AI): Building MLOps and Kubernetes-based platform infrastructure for a large-scale, energy-efficient AI cloud with an accent on infrastructure automation, observability, security, and NVIDIA GPU networking. Focus on scaling platform engineering capabilities, integrating NVIDIA software services, developing self-service tools, and leading incident response and root-cause analysis.

Location: Singapore

Company

hirify.global Technologies develops and operates energy-efficient AI infrastructure, including the hirify.global AI Cloud GPU platform, by combining AI software orchestration, liquid cooling, energy management, and construction.

What you will do

  • Build MLOps capabilities for reproducible, scalable, and secure machine-learning workflows across internal and customer-facing environments.
  • Improve the DevOps platform, including CI/CD and infrastructure-service integrations, for reliability, scalability, and security.
  • Design, implement, operate, and secure Kubernetes production infrastructure supporting NVIDIA GB300 NVL72 systems with Quantum-X800 InfiniBand or Spectrum-X Ethernet.
  • Develop and scale observability platforms using telemetry, OpenTelemetry, and related monitoring technologies.
  • Integrate central services with NVIDIA Mission Control, NETQ, UFM, and NMX, while enhancing internal self-service platform products.
  • Lead incident response, participate in on-call rotation, and conduct root-cause analyses to improve operational maturity.

Requirements

  • Bachelor’s degree in computer science or a related technical field.
  • 7+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer, MLOps Engineer, or Observability Engineer.
  • Strong experience with infrastructure-as-code, configuration management, CI/CD, Terraform, Ansible, GitHub Actions, Jenkins, or ArgoCD.
  • Experience with Docker, Kubernetes networking and cluster management, observability stacks, telemetry solutions, OpenTelemetry, and compliance automation.
  • Competent scripting and programming skills in Bash, Python, or Go, plus knowledge of Linux internals, networking stacks, and distributed storage.
  • Clear and effective English communication, written and spoken.

Nice to have

  • Experience in high-growth startups or regulated industries with strong security and data-privacy requirements.
  • Experience with SOC 2 Type 2 and ISO 27001.

Culture & Benefits

  • Work at the intersection of sustainability, artificial intelligence, and next-generation infrastructure.
  • Collaborate closely with founders and experienced engineers in an emerging company.
  • Contribute to making AI compute more sustainable, accessible, and affordable.
  • Inclusive environment focused on diverse backgrounds and authentic collaboration.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’