Назад
1 месяц назад

Principal Solution Specialist, Core Services (AI Infrastructure)

198 000 - 264 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Solution Specialist, Core Services (AI Infrastructure) (GPU compute, networking, and Kubernetes): Driving new customer wins for AI infrastructure by connecting GPU performance, cluster networking, and platform reliability to enterprise and research workloads with an accent on GPU cluster architecture, RDMA fabrics, and large-scale capacity commitments. Focus on designing reservation deals, translating customer requirements into compute and networking roadmap input, and solving performance and scalability challenges across distributed training and inference deployments.

Location: San Francisco, California or Seattle, Washington. Applicants must meet U.S. export-control eligibility requirements, such as being a U.S. person or eligible to access the controlled information or obtain the required authorization.

Salary: $198,000–$264,000 base salary per year, plus potential bonus, equity awards, and benefits.

Company

CoreWeave provides cloud infrastructure for building and scaling AI workloads, including GPU compute, high-performance networking, and Kubernetes services.

What you will do

  • Own commercial and technical strategies for new AI infrastructure customer wins.
  • Develop opportunities involving GPU performance, cluster networking, bare-metal infrastructure, and platform reliability.
  • Translate requirements for GPU topology, RDMA networking, and Kubernetes scheduling into product roadmap feedback.
  • Create deal structures, technical playbooks, benchmarks, and sales enablement materials.
  • Advise enterprise and research customers on GPU cluster architecture, network fabrics, instance types, and topology trade-offs.
  • Design large-scale GPU reservation frameworks covering cluster sizing, MFU modeling, and network bandwidth commitments.

Requirements

  • 10+ years of experience in HPC, data center infrastructure, or GPU cluster engineering with customer or revenue impact.
  • 5+ years working with large-scale NVIDIA or AMD GPU clusters, RDMA networking, and distributed training frameworks.
  • Deep knowledge of InfiniBand, RoCE, NVLink, network topology, bandwidth, and latency.
  • Experience deploying and tuning Kubernetes clusters for GPU workloads, including GPU Operator, MIG, and time-slicing.
  • Understanding of compute instance types, CPU-GPU memory bandwidth, NVMe storage, power, cooling, BGP, peering, and data center interconnects.
  • Familiarity with PyTorch FSDP, Megatron-LM, DeepSpeed, MFU, and the performance and fault-tolerance implications of cluster architecture.

Nice to have

  • Experience with large-scale AI training, genomics, financial risk modeling, or government and defense HPC.
  • Background in technical sales, solution consulting, or product management for GPU clusters or data center infrastructure.
  • Understanding of GPU procurement economics and reserved versus on-demand capacity.
  • Advanced degree in Computer Science, Electrical Engineering, or a related field.

Culture & Benefits

  • Entrepreneurial, collaborative environment focused on independent thinking and innovative solutions.
  • Medical, dental, and vision insurance fully paid by CoreWeave, plus life and disability insurance.
  • 401(k) with employer match, HSA, FSA, tuition reimbursement, and employee stock purchase participation.
  • Paid parental leave, family-forming support, childcare support, mental wellness benefits, and flexible PTO.
  • Office and data center amenities include catered lunches and a casual work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →