Назад
обновлено 4 дня назад

Platform Infrastructure Engineer (AI)

20 833 - 40 417$
Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Platform Infrastructure Engineer (AI): Build and operate a unified self-serve compute platform for GPU training and inference workloads with an accent on multi-cloud GPU orchestration, scheduling, and fault tolerance. Focus on designing scalable Kubernetes operators, managing GPU clusters at scale, and ensuring high availability for diverse AI workloads.

Location: San Francisco, New York City, Seattle, or Remote (United States only)

Salary: $250K – $485K (USD, U.S.-based positions)

Company

Perplexity serves hundreds of millions of AI queries monthly, operating a large GPU fleet across multiple cloud providers to support inference and training workloads.

What you will do

  • Build and own a self-serve compute platform enabling inference engineers and researchers to run training and inference workloads without managing infrastructure details.
  • Operate and manage GPU fleet provisioning, lifecycle, reliability, and capacity integration across multiple cloud providers.
  • Develop scheduling and placement logic to efficiently allocate GPU resources under real constraints.
  • Support both long-running distributed training jobs and low-latency production inference services on the same infrastructure.
  • Develop Kubernetes operators and CRDs for GPU orchestration across multi-cloud clusters.
  • Implement fault tolerance, autoscaling, and observability to maintain high utilization and workload resilience.

Requirements

  • Location: Must be based in the United States or able to work remotely within the United States
  • Deep Kubernetes experience including custom operators, CRDs, and multi-cluster federation.
  • Experience managing large-scale GPU clusters with NVIDIA hardware, CUDA, and high-speed networking (InfiniBand or RoCE).
  • Proven ability to orchestrate compute resources across multiple cloud providers (e.g., AWS, GCP, CoreWeave).
  • Strong fundamentals in distributed systems: scheduling, resource allocation, and fault tolerance.
  • Proficiency in infrastructure and systems-level programming languages such as Go, Rust, or C++.

Nice to have

  • Experience with inference serving stacks like vLLM, SGLang, or TensorRT-LLM.
  • Familiarity with HPC schedulers such as Slurm.
  • GPU kernel development experience in CUDA or Triton.
  • Knowledge of high-speed interconnects like InfiniBand, RoCE, or RDMA in production environments.
  • Experience with observability tools for ML workloads such as Prometheus, Grafana, or Weights & Biases.

Culture & Benefits

  • Comprehensive benefits for full-time U.S. employees including equity, health, dental, vision, retirement, fitness, commuter, and dependent care accounts.
  • International full-time employees receive benefits tailored to their region.
  • Work in a high-impact AI infrastructure team supporting hundreds of millions of queries monthly.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →