Назад
3 дня назад

Principal Infrastructure Engineer, AI Cluster Performance & Validation

Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Principal Infrastructure Engineer, AI Cluster Performance & Validation (AI/HPC infrastructure): Building control-plane systems, validation pipelines, and performance tooling for multi-thousand-GPU AI and HPC clusters with an accent on distributed workloads, cluster qualification, and full-stack performance diagnosis. Focus on isolating GPU, networking, storage, scheduler, and communication-library bottlenecks while automating burn-in, telemetry correlation, regression detection, and fleet-wide optimization.

Location: Houston, New York, San Francisco, or Seattle, United States

Company

Nscale operates AI infrastructure and large-scale AI and high-performance computing environments.

What you will do

  • Define architecture, roadmaps, acceptance criteria, and performance standards for multi-thousand-GPU clusters.
  • Run distributed training and inference workloads across large-scale AI clusters to validate production behavior.
  • Lead root-cause analysis of cluster failures and performance regressions across GPUs, networking, storage, schedulers, and communication libraries.
  • Design automated validation, burn-in, benchmarking, thermal, power, and reliability qualification systems for nodes, racks, and full pods.
  • Optimize fabric configuration, collective communication, RDMA and storage paths, GPU settings, NUMA placement, and host performance.
  • Build production-grade Python tooling for automated triage, telemetry correlation, regression detection, and operational workflows.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
  • 10+ years of experience building, operating, or debugging large-scale compute infrastructure, including significant staff- or principal-level technical leadership.
  • Hands-on experience running distributed AI training or inference at scale with parallelism strategies and frameworks such as PyTorch, Megatron-LM, or DeepSpeed.
  • Experience validating and accepting clusters with thousands of GPUs and defending benchmark results to engineers and customers.
  • Deep expertise in Linux systems, Python, high-performance networking, InfiniBand or RoCEv2, RDMA, GPUDirect, NUMA, PCIe, and performance debugging.
  • Experience operating AI workloads under SLURM or Kubernetes at scale, plus experience with C/C++ or Go and configuration management tools such as Ansible or Terraform.

Nice to have

  • Master’s degree or PhD in Engineering, Computer Science, or a related field.
  • Experience bringing up a greenfield GPU supercluster and standardizing firmware, drivers, and topology across a heterogeneous fleet.
  • Experience with NVIDIA GPU platforms, NVLink, NVSwitch, DCGM, SHARP, UFM, or AMD Instinct and ROCm/RCCL.
  • Experience with MLPerf, advanced GPU and fabric observability, parallel storage, cloud-native infrastructure, and bare-metal tooling.
  • Experience using AI tools to improve infrastructure workflows and operator experience.

Culture & Benefits

  • Cross-functional collaboration with Infrastructure, Platform, SRE, and customer-facing teams.
  • Responsibility for engineering standards covering reliability, observability, benchmarking, and operational excellence.
  • Opportunity to mentor engineers, create runbooks, and produce post-incident technical write-ups.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’