3 дня назад
AI Platform Support Engineer (APAC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Platform Support Engineer (APAC) (Kubernetes/GPU/ML infrastructure): Supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms with an accent on distributed systems troubleshooting, platform reliability, and customer-facing technical guidance. Focus on diagnosing PyTorch, CUDA, NCCL, networking, storage, and inference performance issues, while building automation, observability, and operational improvements.
Location: Remote; candidates must be based in the Philippines or Singapore. Sunday-Wednesday shift, 7:00 AM to 5:00 PM local time (UTC+8).
Company
builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-first software with large-scale AI compute infrastructure.
What you will do
- Partner directly with customer ML engineering teams running production training and inference workloads.
- Diagnose distributed systems and ML infrastructure issues across Kubernetes, cloud infrastructure, GPU platforms, and bare-metal environments.
- Troubleshoot PyTorch, CUDA, NCCL, inference serving, networking, storage, and performance-related problems.
- Analyze logs, metrics, traces, and system behavior to identify root causes and resolve high-impact incidents.
- Drive reliability improvements through post-incident reviews, observability enhancements, tooling, automation, documentation, and runbooks.
- Collaborate with infrastructure, networking, and platform engineering teams to improve customer troubleshooting workflows.
Requirements
- Strong software engineering and systems troubleshooting background.
- Experience with Kubernetes, containerized environments, cloud infrastructure, and distributed systems.
- Strong Linux knowledge, including networking, storage, process management, and performance tuning.
- Production or research experience operating machine learning workloads, including distributed ML systems such as PyTorch, CUDA, or NCCL.
- Experience with GPU infrastructure, orchestration, observability, and troubleshooting ML performance, reliability, or scaling issues.
- Strong communication skills and ability to work directly with technical customers and engineering teams in ambiguous environments.
Nice to have
- Experience with large-scale model training, distributed inference, Ray, Kubeflow, Slurm, or similar scheduling platforms.
- Experience with InfiniBand, RDMA, high-performance networking, bare-metal infrastructure, or ML storage systems.
- Experience at an AI infrastructure, cloud, MLOps, or developer tooling company.
- Experience writing Python automation, tooling, or scripts.
Culture & Benefits
- Builder-oriented environment focused on ownership, open communication, continuous improvement, and long-term thinking.
- Medical, dental, and vision coverage for employees and eligible dependents.
- Equity, location-dependent retirement or pension contributions, unlimited PTO, company holidays, and a two-week winter break.
- Paid parental and family leave, professional development allowance, wellness benefits, and work-from-home stipends.
- Four weeks of paid sabbatical leave after four years of service.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →