Назад
Company hidden
5 часов назад

Supercomputing Platform & Infrastructure Engineer (AI)

200 000 - 550 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Релокация
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Supercomputing Platform & Infrastructure Engineer (AI): Designing and operating large-scale GPU clusters for model training and inference with an accent on Terraform-based infrastructure-as-code, Kubernetes orchestration, and high-throughput networking and storage. Focus on automating fault detection and recovery, debugging cross-layer failures, and improving observability and reliability across distributed infrastructure.

Location: On-site in San Francisco, with relocation support to SF where possible

Salary: $200,000–$550,000 annual base salary, depending on experience; equity is also included in total compensation.

Company

hirify.global is building safe artificial general intelligence through research automation, code generation, frontier-scale pre-training, domain-specific reinforcement learning, long-context models, and inference-time compute.

What you will do

  • Design, operate, and scale GPU clusters supporting model training and inference workloads.
  • Build and maintain Terraform-driven infrastructure across cloud and hybrid environments.
  • Deploy, operate, and optimize Kubernetes clusters for AI workload scheduling.
  • Develop scalable infrastructure-as-code patterns for compute, networking, and storage provisioning.
  • Optimize high-throughput networking and storage, deployment reproducibility, and environment consistency.
  • Automate fault detection and recovery while improving observability and platform reliability.

Requirements

  • Strong software engineering skills and experience building production infrastructure systems.
  • Deep hands-on Terraform experience, including module design, state management, environment isolation, and large-scale deployments.
  • Experience operating production GPU infrastructure or high-performance distributed systems.
  • Strong understanding of networking and storage systems.
  • Experience with major cloud platforms such as GCP, AWS, Azure, or OCI.
  • Experience owning production-critical infrastructure end to end and debugging issues across hardware, drivers, networking, storage, operating systems, and cloud layers.

Culture & Benefits

  • Small, fast-paced, highly focused team working toward safe AGI.
  • Equity as a significant part of total compensation.
  • 401(k) plan with 6% salary matching.
  • Health, dental, and vision insurance for employees and dependents.
  • Unlimited paid time off.
  • Visa sponsorship and relocation stipend to San Francisco where possible.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →