Назад
Company hidden
2 дня назад

Forward Deployed Engineer (AI Infrastructure)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Forward Deployed Engineer (AI Infrastructure): Building and optimizing large-scale GPU clusters for AI training and inference with an accent on reliability, interconnect performance, and customer onboarding. Focus on debugging distributed training failures, tuning network fabrics (InfiniBand/RoCE), and automating cluster health management.

Location: Remote (North America) / San Francisco, CA

Company

hirify.global provides early-stage startups with access to scaled AI infrastructure and compute liquidity for training and inference.

What you will do

  • Serve as the primary technical contact for teams running large-scale training and inference workloads, owning onboarding and environment setup.
  • Diagnose and fix real-world failures including NCCL timeouts, OOM patterns, and driver mismatches within customer environments.
  • Profile and improve distributed training performance to reduce GPU idle time and improve MFU.
  • Ensure the health of high-speed interconnects (InfiniBand, RoCE, NVLink) and resolve fabric-level issues.
  • Develop automation for cluster provisioning, GPU health checks, and firmware/driver lifecycle management.
  • Lead incident response and communication for complex failures spanning hardware, networking, and ML frameworks.

Requirements

  • Hands-on experience operating production GPU clusters (NVIDIA A100/H100/B200).
  • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training.
  • Expert-level Linux knowledge, including kernel tuning, CUDA toolkit, and performance profiling.
  • Strong experience running Kubernetes for GPU workloads or HPC schedulers like Slurm.
  • Engineering proficiency in Python, Go, or Bash to build production-grade tools.
  • Must be based in North America.

Nice to have

  • Experience with high-performance parallel file systems (VAST, WEKA, Lustre, GPFS).
  • Background in solutions architecture or technical account ownership at an infrastructure company.
  • Hands-on experience with production inference, including autoscaling and KV cache behavior.
  • Contributions to relevant OSS projects or published deep-dives on distributed systems.

Culture & Benefits

  • High-growth environment at the center of the AI infrastructure boom.
  • High level of ownership as the first FDE on the solutions engineering team.
  • Competitive compensation with meaningful equity.
  • Comprehensive healthcare, dental, and vision coverage for employees and dependents.
  • 401(k) and unlimited PTO.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →