HPC Architect (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
HPC Architect (AI Infrastructure): Establishing technical qualification standards and vetting global compute providers with an accent on cluster architecture, network fabric, and hardware performance. Focus on building a rigorous acceptance test suite and validating GPU clusters to ensure production-grade stability.
Location: North America Remote / San Francisco, CA
Company
provides early-stage startups with access to scaled AI infrastructure, building a liquidity layer for global AI compute.
What you will do
- Vet prospective compute providers by assessing cluster architecture, GPU hardware, network fabric, storage, and orchestration.
- Define the technical qualification bar, building acceptance test suites and benchmark methodologies from scratch.
- Perform hands-on validation, including burn-in testing, fabric validation (InfiniBand/RoCE), and NCCL benchmarks.
- Guide provider data-center and platform engineers through technical onboarding and configuration tuning.
- Partner with the compute procurement team to perform technical due diligence on new providers.
- Maintain technical relationships with existing providers to identify architectural or quality drift.
Requirements
- Deep HPC experience designing, building, or operating GPU clusters at scale.
- Strong knowledge of network fabrics, specifically InfiniBand and RoCE.
- Experience with distributed orchestration and HPC software stacks (Slurm, Kubernetes, OpenMPI, Linux).
- Data-center literacy covering power, cooling, cabling, and physical-layer realities.
- Ability to design non-gameable benchmarking tests for large-scale training workloads.
- Capability to write precise, testable technical standards for external engineering teams.
Nice to have
- Experience working within a neocloud, hyperscaler, or colocation provider.
- Deep expertise in NVIDIA data-center GPU platforms (DGX/HGX, NVLink/NVSwitch).
- Experience supporting AI research labs and familiarity with training workload bottlenecks.
- Background in cluster acceptance testing or site bring-up.
Culture & Benefits
- High-growth environment at the center of the AI infrastructure boom.
- High level of ownership as the first HPC Architect for the solutions engineering team.
- Competitive compensation with meaningful equity.
- Comprehensive healthcare, dental, and vision coverage for employees and dependents.
- 401(k) retirement plan and unlimited PTO.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →