Назад
Company hidden
5 часов назад

Infrastructure Platform Engineer (AI)

150 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Infrastructure Platform Engineer (AI): Building and operating heterogeneous CPU, GPU, and accelerator clusters for production AI inference with an accent on Linux systems, Kubernetes, provisioning automation, scheduling, networking, and observability. Focus on bringing new accelerator platforms into production, managing cluster capacity and isolation, and solving complex reliability and performance challenges across compute, memory, network, and storage.

Location: San Francisco, CA; on-site

Salary: $150,000–$350,000 per year, plus equity

Company

hirify.global is building a multi-silicon neocloud for fast, efficient AI inference across heterogeneous hardware.

What you will do

  • Design, deploy, and operate large-scale CPU, GPU, and accelerator clusters for production AI inference.
  • Build provisioning and lifecycle management systems for deployment, upgrades, validation, and fleet operations.
  • Improve scheduling, resource utilization, isolation, and capacity management across heterogeneous hardware.
  • Build observable infrastructure for debugging, incident response, and reliable production operations.
  • Bring new accelerator platforms into production in partnership with distributed systems, runtime, compiler, networking, and hardware engineers.
  • Influence the architecture of the infrastructure platform supporting next-generation AI workloads.

Requirements

  • Experience in infrastructure, cluster engineering, platform engineering, SRE, HPC, or distributed systems.
  • Deep Linux systems experience, including performance, networking, storage, process, and kernel-level debugging.
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration and scheduling systems.
  • Strong automation skills with Terraform, Ansible, Helm, Python, Go, or equivalent tools.
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA/ROCm stacks, or hardware validation.
  • Familiarity with high-performance networking and the ability to build observable, recoverable production systems.

Nice to have

  • Experience with AI inference, training, HPC, or neocloud infrastructure.
  • Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up.
  • Experience with multi-tenant isolation, quota systems, fair scheduling, or usage accounting.
  • Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks.
  • Experience with observability platforms such as Prometheus, OpenTelemetry, or Grafana, and heterogeneous hardware from NVIDIA, AMD, Intel, ARM, or emerging accelerator vendors.

Culture & Benefits

  • Early-stage startup environment with significant ownership and ambiguity.
  • Close collaboration with a small group of highly capable engineers.
  • Opportunity to shape the infrastructure platform and how the company scales.
  • Equity is offered as part of compensation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →