3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Principal Infrastructure Engineer, AI Cluster Performance & Validation
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Principal Infrastructure Engineer, AI Cluster Performance & Validation (AI/HPC infrastructure): Building control-plane systems, validation pipelines, and performance tooling for multi-thousand-GPU AI and HPC clusters with an accent on distributed workloads, cluster qualification, and full-stack performance diagnosis. Focus on isolating GPU, networking, storage, scheduler, and communication-library bottlenecks while automating burn-in, telemetry correlation, regression detection, and fleet-wide optimization.
Location: Houston, New York, San Francisco, or Seattle, United States
Company
Nscale operates AI infrastructure and large-scale AI and high-performance computing environments.
What you will do
- Define architecture, roadmaps, acceptance criteria, and performance standards for multi-thousand-GPU clusters.
- Run distributed training and inference workloads across large-scale AI clusters to validate production behavior.
- Lead root-cause analysis of cluster failures and performance regressions across GPUs, networking, storage, schedulers, and communication libraries.
- Design automated validation, burn-in, benchmarking, thermal, power, and reliability qualification systems for nodes, racks, and full pods.
- Optimize fabric configuration, collective communication, RDMA and storage paths, GPU settings, NUMA placement, and host performance.
- Build production-grade Python tooling for automated triage, telemetry correlation, regression detection, and operational workflows.
Requirements
- Bachelorβs degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
- 10+ years of experience building, operating, or debugging large-scale compute infrastructure, including significant staff- or principal-level technical leadership.
- Hands-on experience running distributed AI training or inference at scale with parallelism strategies and frameworks such as PyTorch, Megatron-LM, or DeepSpeed.
- Experience validating and accepting clusters with thousands of GPUs and defending benchmark results to engineers and customers.
- Deep expertise in Linux systems, Python, high-performance networking, InfiniBand or RoCEv2, RDMA, GPUDirect, NUMA, PCIe, and performance debugging.
- Experience operating AI workloads under SLURM or Kubernetes at scale, plus experience with C/C++ or Go and configuration management tools such as Ansible or Terraform.
Nice to have
- Masterβs degree or PhD in Engineering, Computer Science, or a related field.
- Experience bringing up a greenfield GPU supercluster and standardizing firmware, drivers, and topology across a heterogeneous fleet.
- Experience with NVIDIA GPU platforms, NVLink, NVSwitch, DCGM, SHARP, UFM, or AMD Instinct and ROCm/RCCL.
- Experience with MLPerf, advanced GPU and fabric observability, parallel storage, cloud-native infrastructure, and bare-metal tooling.
- Experience using AI tools to improve infrastructure workflows and operator experience.
Culture & Benefits
- Cross-functional collaboration with Infrastructure, Platform, SRE, and customer-facing teams.
- Responsibility for engineering standards covering reliability, observability, benchmarking, and operational excellence.
- Opportunity to mentor engineers, create runbooks, and produce post-incident technical write-ups.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
HPC Engineer (AI Cloud)
6 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Forward Deployed Engineer (AI Infrastructure)
4 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Senior HPC Engineer (Classified Computing)
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Intelligent Data Infrastructure Engineer (AI)
196Β 350 - 292Β 600$
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Cloud Networking & Infrastructure Developer (AWS)
6 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Principal Computational Engineer (HPC)
114Β 400 - 216Β 320$