Назад
Company hidden
3 дня назад

AI Platform Support Engineer (APAC)

Формат работы
remote (только Philippines/Singapore)
Тип работы
fulltime
Английский
b2
Страна
Singapore/Philippines
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Platform Support Engineer (APAC) (Kubernetes/GPU/ML infrastructure): Supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms with an accent on distributed systems troubleshooting, platform reliability, and customer-facing technical guidance. Focus on diagnosing PyTorch, CUDA, NCCL, networking, storage, and inference performance issues, while building automation, observability, and operational improvements.

Location: Remote; candidates must be based in the Philippines or Singapore. Sunday-Wednesday shift, 7:00 AM to 5:00 PM local time (UTC+8).

Company

hirify.global builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-first software with large-scale AI compute infrastructure.

What you will do

  • Partner directly with customer ML engineering teams running production training and inference workloads.
  • Diagnose distributed systems and ML infrastructure issues across Kubernetes, cloud infrastructure, GPU platforms, and bare-metal environments.
  • Troubleshoot PyTorch, CUDA, NCCL, inference serving, networking, storage, and performance-related problems.
  • Analyze logs, metrics, traces, and system behavior to identify root causes and resolve high-impact incidents.
  • Drive reliability improvements through post-incident reviews, observability enhancements, tooling, automation, documentation, and runbooks.
  • Collaborate with infrastructure, networking, and platform engineering teams to improve customer troubleshooting workflows.

Requirements

  • Strong software engineering and systems troubleshooting background.
  • Experience with Kubernetes, containerized environments, cloud infrastructure, and distributed systems.
  • Strong Linux knowledge, including networking, storage, process management, and performance tuning.
  • Production or research experience operating machine learning workloads, including distributed ML systems such as PyTorch, CUDA, or NCCL.
  • Experience with GPU infrastructure, orchestration, observability, and troubleshooting ML performance, reliability, or scaling issues.
  • Strong communication skills and ability to work directly with technical customers and engineering teams in ambiguous environments.

Nice to have

  • Experience with large-scale model training, distributed inference, Ray, Kubeflow, Slurm, or similar scheduling platforms.
  • Experience with InfiniBand, RDMA, high-performance networking, bare-metal infrastructure, or ML storage systems.
  • Experience at an AI infrastructure, cloud, MLOps, or developer tooling company.
  • Experience writing Python automation, tooling, or scripts.

Culture & Benefits

  • Builder-oriented environment focused on ownership, open communication, continuous improvement, and long-term thinking.
  • Medical, dental, and vision coverage for employees and eligible dependents.
  • Equity, location-dependent retirement or pension contributions, unlimited PTO, company holidays, and a two-week winter break.
  • Paid parental and family leave, professional development allowance, wellness benefits, and work-from-home stipends.
  • Four weeks of paid sabbatical leave after four years of service.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →