Forward Deployed Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Forward Deployed Engineer (AI Infrastructure): Building and optimizing large-scale GPU clusters for AI training and inference with an accent on reliability, interconnect performance, and customer onboarding. Focus on debugging distributed training failures, tuning network fabrics (InfiniBand/RoCE), and automating cluster health management.
Location: Remote (North America) / San Francisco, CA
Company
provides early-stage startups with access to scaled AI infrastructure and compute liquidity for training and inference.
What you will do
- Serve as the primary technical contact for teams running large-scale training and inference workloads, owning onboarding and environment setup.
- Diagnose and fix real-world failures including NCCL timeouts, OOM patterns, and driver mismatches within customer environments.
- Profile and improve distributed training performance to reduce GPU idle time and improve MFU.
- Ensure the health of high-speed interconnects (InfiniBand, RoCE, NVLink) and resolve fabric-level issues.
- Develop automation for cluster provisioning, GPU health checks, and firmware/driver lifecycle management.
- Lead incident response and communication for complex failures spanning hardware, networking, and ML frameworks.
Requirements
- Hands-on experience operating production GPU clusters (NVIDIA A100/H100/B200).
- Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training.
- Expert-level Linux knowledge, including kernel tuning, CUDA toolkit, and performance profiling.
- Strong experience running Kubernetes for GPU workloads or HPC schedulers like Slurm.
- Engineering proficiency in Python, Go, or Bash to build production-grade tools.
- Must be based in North America.
Nice to have
- Experience with high-performance parallel file systems (VAST, WEKA, Lustre, GPFS).
- Background in solutions architecture or technical account ownership at an infrastructure company.
- Hands-on experience with production inference, including autoscaling and KV cache behavior.
- Contributions to relevant OSS projects or published deep-dives on distributed systems.
Culture & Benefits
- High-growth environment at the center of the AI infrastructure boom.
- High level of ownership as the first FDE on the solutions engineering team.
- Competitive compensation with meaningful equity.
- Comprehensive healthcare, dental, and vision coverage for employees and dependents.
- 401(k) and unlimited PTO.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →