Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
GPU Cluster Architect (AI): Designing next-generation GPU cluster infrastructure across compute, networking, storage, and control planes with an accent on massive-scale AI workloads, performance, and reliability. Focus on modeling LLM training and inference requirements, validating InfiniBand and RoCEv2 interconnects, optimizing storage integration, and detecting design flaws through monitoring signals.
Location: Remote - United States. Applicants must be authorized to work in the country where they apply.
Salary: $184K–$318K OTE, including base salary and performance bonus. RSUs may be available at certain salary grades.
Company
Nebius builds a full-stack AI cloud platform for data processing, model training, and production deployment, with infrastructure spanning compute, storage, networking, and applied AI.
What you will do
- Architect scalable GPU cluster topologies covering compute nodes, InfiniBand or Ethernet interconnects, storage, and control planes.
- Model LLM training and inference workloads to evaluate latency, bandwidth, GPU density, and other architectural trade-offs.
- Design and validate low-latency, high-throughput interconnects including InfiniBand HDR/NDR and RoCEv2 at pod and data-center scale.
- Integrate storage for training datasets, checkpointing, and related workloads.
- Analyze monitoring signals to identify design issues and improve reliability.
- Collaborate with site reliability, networking, storage, and data-center engineering teams to operationalize and scale the architecture.
Requirements
- 5+ years of experience designing clusters.
- Deep understanding of modern GPU architectures, including NVIDIA and AMD.
- Experience with HPC interconnects, including InfiniBand and RoCE.
- Strong background in systems architecture, networking, and hardware reliability.
- Experience scripting automation and telemetry pipelines with Python, Go, or similar technologies.
- Authorization to work in the United States is required.
Culture & Benefits
- Remote work with up to $85 per month reimbursement for mobile and internet expenses.
- Company-paid medical, dental, and vision insurance for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- Paid parental leave, including 20 weeks for primary caregivers and 12 weeks for secondary caregivers.
- Company-paid short-term, long-term, and life insurance.
- Opportunities for professional growth, learning, ownership, and collaboration on AI infrastructure projects.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Senior Solutions Architect (AI, HPC & Lustre)
197 200 - 255 200$
Lambda
9 дней назад
Data Center Standards Architect (AI)
443 000 - 590 000$
Baseten
5 дней назад
Solution Architect (AI)
165 000 - 330 000$
Anthropic
5 дней назад
Applied AI Architect (Enterprise Tech)
240 000 - 315 000$
7 дней назад
AI Data Platform Field Architect
153 500 - 298 500$
10 дней назад
Cloud-Native Architect (Kubernetes)
130 000 - 180 000$