Назад
Company hidden
6 часов назад

Senior Infrastructure Engineer (AI Cloud)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Infrastructure Engineer (AI Cloud): Architecting and building a managed AI cloud platform spanning GPU capacity, control-plane services, customer interfaces, storage, and workload orchestration with an accent on multi-tenancy, automation, and high-performance infrastructure. Focus on designing production cloud services, integrating heterogeneous GPU providers, operating Kubernetes and Slurm, and delivering reliable metering, billing, and observability.

Location: Hybrid in the Bay Area, Boston, or Washington D.C.; 2 work-from-home days per week

Company

hirify.global develops software that makes AI data centers power-flexible, helping scale compute while supporting grid reliability and renewable energy growth.

What you will do

  • Architect and launch managed cloud services for GPU capacity, including tenant models, isolation boundaries, provisioning flows, and service catalogs.
  • Build control-plane services, self-service customer interfaces, lifecycle automation, usage metering, and billing integrations.
  • Assess and integrate bare-metal GPU infrastructure providers while designing portability across vendors.
  • Implement secure multi-tenancy across compute, storage, and networking, including InfiniBand/VLAN isolation, QoS, encryption, and root-access boundaries.
  • Operate Kubernetes and Slurm environments for large-scale training and inference across heterogeneous clouds.
  • Define storage, SLO, observability, and incident-response strategies for a reliable customer-facing service.

Requirements

  • 7+ years of infrastructure or platform engineering experience, including launching a managed cloud or AI platform used by production customers.
  • Strong production experience with Kubernetes and Slurm as managed services.
  • Production experience with Lustre or comparable parallel filesystems such as GPFS, Weka, VAST, or BeeGFS.
  • Knowledge of cloud control planes, APIs, tenancy and isolation models, quotas, metering, and customer-facing operations.
  • Deep Linux expertise, infrastructure as code with Terraform or Ansible, and programming ability in Python or Go.
  • Experience with GPU infrastructure, InfiniBand, RoCE, RDMA, and the GPU software stack.

Nice to have

  • Experience with NVIDIA SuperPOD, GPUDirect Storage, NCCL debugging, or DCGM.
  • Experience with Lustre multitenancy, VAST or Weka service-provider deployments, and object storage such as S3, Ceph, or MinIO.
  • Experience negotiating with infrastructure vendors and building billing, metering, or FinOps pipelines.

Culture & Benefits

  • Collaborative, low-ego environment with AI, cloud, software, and energy specialists.
  • Opportunity to shape strategy, go-to-market, organizational design, and customer engagement from the beginning.
  • Competitive pay, equity, medical, dental, and vision coverage.
  • 401(k) matching and a flexible hybrid work arrangement.
  • Equal opportunity workplace with reasonable accommodations for applicants with disabilities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →