Назад
Company hidden
3 дня назад

AI Engineer, AI & Applications

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Engineer, AI & Applications (Distributed AI Training): Building production-ready training recipes, benchmarking suites, and evaluation harnesses for efficient distributed model training with an accent on GPU optimization, model parallelism, and rigorous performance measurement. Focus on designing multi-node NCCL workflows, optimizing PyTorch/JAX workloads at scale, and creating actionable playbooks for hyperscale AI customers.

Location: Singapore or Australia, specifically Launceston, Tasmania or Sydney, New South Wales

Company

hirify.global develops sustainable AI infrastructure and production-grade systems for hyperscale model training and AI applications.

What you will do

  • Build production-ready TorchTitan and Megatron-LM training recipes, including model configurations, parallelism strategies, and checkpointing patterns.
  • Document tuning parameters and expected throughput for distributed training workloads at different scales.
  • Create and validate multi-node NCCL communication patterns on AI Factory Kubernetes and Slurm clusters.
  • Design benchmarking suites covering accuracy, latency, throughput, cost per token, energy efficiency, and MFU.
  • Implement offline evaluation harnesses, leaderboard tracking, and fine-tuning experiments using LoRA and QLoRA.
  • Partner with scheduling, orchestration, AI, and software engineers to integrate templates and optimize models for inference and AI applications.

Requirements

  • 5–7 years of experience in distributed machine learning with PyTorch or JAX, including multi-node training across 10+ GPUs.
  • Expert knowledge of GPU utilization, memory behavior, NCCL collectives, communication bottlenecks, and throughput optimization.
  • Hands-on experience debugging convergence issues, profiling bottlenecks, and optimizing distributed training at scale.
  • Strong benchmarking methodology, including controlled experiments, noise measurement, variance analysis, and rigorous communication of results.
  • Familiarity with TorchTitan, Megatron-LM, or comparable production training frameworks.
  • Understanding of FSDP, tensor parallelism, pipeline parallelism, checkpointing, recovery, resource constraints, and cost optimization.

Culture & Benefits

  • Full-time employment.
  • Inclusive workplace welcoming candidates from diverse backgrounds.
  • Opportunity to contribute to sustainable engineering practices in AI infrastructure.
  • Reporting to the Head of AI & Applications.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →