Назад
обновлено 4 дня назад

GPU Cluster Architect (AI)

184 000 - 318 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
GPU Cluster Architect (AI): Designing next-generation GPU cluster infrastructure across compute, networking, storage, and control planes with an accent on massive-scale AI workloads, performance, and reliability. Focus on modeling LLM training and inference requirements, validating InfiniBand and RoCEv2 interconnects, optimizing storage integration, and detecting design flaws through monitoring signals.

Location: Remote - United States. Applicants must be authorized to work in the country where they apply.

Salary: $184K–$318K OTE, including base salary and performance bonus. RSUs may be available at certain salary grades.

Company

Nebius builds a full-stack AI cloud platform for data processing, model training, and production deployment, with infrastructure spanning compute, storage, networking, and applied AI.

What you will do

  • Architect scalable GPU cluster topologies covering compute nodes, InfiniBand or Ethernet interconnects, storage, and control planes.
  • Model LLM training and inference workloads to evaluate latency, bandwidth, GPU density, and other architectural trade-offs.
  • Design and validate low-latency, high-throughput interconnects including InfiniBand HDR/NDR and RoCEv2 at pod and data-center scale.
  • Integrate storage for training datasets, checkpointing, and related workloads.
  • Analyze monitoring signals to identify design issues and improve reliability.
  • Collaborate with site reliability, networking, storage, and data-center engineering teams to operationalize and scale the architecture.

Requirements

  • 5+ years of experience designing clusters.
  • Deep understanding of modern GPU architectures, including NVIDIA and AMD.
  • Experience with HPC interconnects, including InfiniBand and RoCE.
  • Strong background in systems architecture, networking, and hardware reliability.
  • Experience scripting automation and telemetry pipelines with Python, Go, or similar technologies.
  • Authorization to work in the United States is required.

Culture & Benefits

  • Remote work with up to $85 per month reimbursement for mobile and internet expenses.
  • Company-paid medical, dental, and vision insurance for employees and families.
  • 401(k) plan with up to a 4% company match and immediate vesting.
  • Paid parental leave, including 20 weeks for primary caregivers and 12 weeks for secondary caregivers.
  • Company-paid short-term, long-term, and life insurance.
  • Opportunities for professional growth, learning, ownership, and collaboration on AI infrastructure projects.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →