3 дня назад
AI Engineer, AI & Applications
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Engineer, AI & Applications (Distributed AI Training): Building production-ready training recipes, benchmarking suites, and evaluation harnesses for efficient distributed model training with an accent on GPU optimization, model parallelism, and rigorous performance measurement. Focus on designing multi-node NCCL workflows, optimizing PyTorch/JAX workloads at scale, and creating actionable playbooks for hyperscale AI customers.
Location: Singapore or Australia, specifically Launceston, Tasmania or Sydney, New South Wales
Company
develops sustainable AI infrastructure and production-grade systems for hyperscale model training and AI applications.
What you will do
- Build production-ready TorchTitan and Megatron-LM training recipes, including model configurations, parallelism strategies, and checkpointing patterns.
- Document tuning parameters and expected throughput for distributed training workloads at different scales.
- Create and validate multi-node NCCL communication patterns on AI Factory Kubernetes and Slurm clusters.
- Design benchmarking suites covering accuracy, latency, throughput, cost per token, energy efficiency, and MFU.
- Implement offline evaluation harnesses, leaderboard tracking, and fine-tuning experiments using LoRA and QLoRA.
- Partner with scheduling, orchestration, AI, and software engineers to integrate templates and optimize models for inference and AI applications.
Requirements
- 5–7 years of experience in distributed machine learning with PyTorch or JAX, including multi-node training across 10+ GPUs.
- Expert knowledge of GPU utilization, memory behavior, NCCL collectives, communication bottlenecks, and throughput optimization.
- Hands-on experience debugging convergence issues, profiling bottlenecks, and optimizing distributed training at scale.
- Strong benchmarking methodology, including controlled experiments, noise measurement, variance analysis, and rigorous communication of results.
- Familiarity with TorchTitan, Megatron-LM, or comparable production training frameworks.
- Understanding of FSDP, tensor parallelism, pipeline parallelism, checkpointing, recovery, resource constraints, and cost optimization.
Culture & Benefits
- Full-time employment.
- Inclusive workplace welcoming candidates from diverse backgrounds.
- Opportunity to contribute to sustainable engineering practices in AI infrastructure.
- Reporting to the Head of AI & Applications.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →