Назад
Company hidden
6 дней назад

Senior Infrastructure Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Infrastructure Engineer (AI): Building and operating large-scale distributed systems for AI cluster management with an accent on Kubernetes operators, bare-metal automation, and observability. Focus on designing high-performance control planes, ensuring system reliability, and scaling infrastructure for wafer-scale AI hardware.

Location: Sunnyvale, CA (Hybrid)

Company

hirify.global Systems builds breakthrough AI hardware, including the world's largest AI chip, to deliver industry-leading training and inference speeds for global enterprises and AI-native startups.

What you will do

  • Develop declarative, CRD-driven automation for bare-metal networking, OS, and application software across large-scale clusters.
  • Build and maintain Kubernetes operators to schedule complex inference workloads with priority queues and health-aware placement.
  • Design gRPC control-plane services, including authorization, admission webhooks, and multi-tenant quota policies.
  • Implement robust metrics and log pipelines using Prometheus and Grafana for wafer-scale systems and network fabric.
  • Ensure system reliability through failure detection, HA control planes, and automated recovery mechanisms.

Requirements

  • 5+ years of experience building and operating production distributed systems or infrastructure software.
  • Proficiency in Go and Python for production-quality code.
  • Deep expertise in Kubernetes, including writing controllers, operators, CRDs, and admission webhooks.
  • Strong debugging skills across distributed systems, Linux, and networking.
  • Practical experience with Prometheus and Grafana, including PromQL and exporter design.
  • Demonstrated adoption of AI tools in engineering workflows with a focus on verification and rigor.

Nice to have

  • Experience with bare-metal or HPC fleet operations.
  • Knowledge of RDMA/RoCE, eBPF, Ceph, NVMe-oF, or etcd.
  • Familiarity with scheduler internals and inference serving stacks.

Culture & Benefits

  • Opportunity to work on one of the fastest AI supercomputers in the world.
  • Non-corporate work culture that respects individual beliefs.
  • Balance of startup vitality with long-term job stability.
  • Support for continuous learning, growth, and open-source research contributions.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →