Назад
Company hidden
4 дня назад

Staff AI Infrastructure Engineer (AI)

241 000 - 331 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff AI Infrastructure Engineer (AI/HPC): Building and operating multi-site GPU clusters running Slurm on Kubernetes to support frontier-scale AI biology research with an accent on reliability, distributed systems, and high-performance computing. Focus on debugging multi-node training failures, scaling storage and networking infrastructure, automating cluster operations, and improving incident resilience across thousands of GPUs.

Location: Redwood City, California, United States; hybrid with onsite attendance at least 60% of the working month, approximately 3 days per week

Salary: $241,000–$331,000 base pay per year

Company

hirify.global is a nonprofit open-science research lab building AI, biological foundation models, and laboratory capabilities to accelerate scientific discovery and help cure disease.

What you will do

  • Own reliability, observability, and incident response for multi-site GPU clusters running Slurm on Kubernetes.
  • Debug infrastructure failures across storage, networking, scheduling, and GPU compute layers.
  • Design and validate cluster scaling plans for larger multi-node and multi-thousand-GPU training runs.
  • Build automation for capacity planning, GPU utilization monitoring, workload-manager policies, pod lifecycle management, and configuration as code.
  • Collaborate with AI researchers and hero-run leads to design infrastructure for frontier-scale workloads.
  • Coordinate technical vendor escalations, improve operational resilience, and develop scalable runbooks.

Requirements

  • 8+ years of AI/ML infrastructure engineering experience with deep expertise in HPC/Slurm operations, Kubernetes at scale, distributed systems debugging, or GPU infrastructure.
  • Strong Linux systems knowledge, including TCP/IP, InfiniBand, RDMA, storage systems, kernel internals, cgroups, namespaces, eBPF, and sysctls.
  • Hands-on Kubernetes and cloud-native experience, including pod lifecycle, CNI plugins, StatefulSets, Helm, and GitOps tooling.
  • Experience with HPC workload managers, especially Slurm, and proficiency in Python and Bash.
  • Experience with observability tools such as Prometheus or VictoriaMetrics, Grafana, DCGM metrics, and distributed tracing.
  • Ability to work onsite in Redwood City at least 60% of the working month.

Nice to have

  • Go, Rust, or C/C++ experience.
  • Experience with Cilium, ArgoCD, or advanced Slurm patterns.
  • Distributed AI training experience with NCCL, PyTorch DDP, multi-node job debugging, and checkpoint/restart workflows.

Culture & Benefits

  • Open-science and open-source AI research focused on improving human health.
  • Collaborative, team-oriented environment with regular in-person work.
  • Employer match on 401(k) contributions.
  • Paid time off for volunteering.
  • Funding for select family-forming benefits and relocation support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →