Назад
Company hidden
10 дней назад

Member of Technical Staff, Infrastructure (AI)

150 000 - 390 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Infrastructure (AI): Building cluster infrastructure for a cloud platform that runs inference workloads across GPUs, CPUs, and emerging accelerator architectures with an accent on bare-metal provisioning, fleet lifecycle management, scheduling, and production reliability. Focus on making new accelerator hardware production-ready, improving observability and recovery, and operating reliable infrastructure at scale.

Location: On-site in San Francisco, California, five days per week

Salary: $150,000–$390,000 base plus equity

Company

Series A AI infrastructure startup building a cloud platform for inference workloads across GPUs, CPUs, and emerging accelerator architectures.

What you will do

  • Deploy production clusters across different accelerator architectures.
  • Automate bare-metal provisioning, validation, upgrades, and fleet lifecycle management.
  • Improve cluster scheduling, resource utilization, isolation, and capacity management.
  • Build observability systems for debugging, incident response, and recovery.
  • Make new accelerators production-ready across drivers, firmware, networking, and orchestration.
  • Partner with runtime, compiler, distributed-systems, networking, and hardware engineers.

Requirements

  • Experience in infrastructure, platform engineering, cluster engineering, SRE, or HPC.
  • Strong Linux systems knowledge and production debugging experience.
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems.
  • Infrastructure automation experience with Python, Go, Terraform, or Ansible.
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA, or ROCm.
  • Experience building observable, recoverable, and reliable production systems at scale.

Culture & Benefits

  • Small, highly technical team.
  • Opportunity to build production infrastructure across multiple generations and types of AI hardware.
  • Equity included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →