Назад
Company hidden
7 часов назад

Senior GPU Systems & Fabric Engineer (AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior GPU Systems & Fabric Engineer (AI): Building the high-performance hardware foundation for AI cloud computing by integrating GPUs, Kubernetes, and low-latency networking with an accent on Linux kernel internals, GPU architectures, and high-speed interconnects. Focus on optimizing RDMA and InfiniBand fabrics, automating GPU and NIC remediation, and enabling topology-aware multi-tenant workloads.

Location: Remote within San Jose, US or Austin, TX

Company

hirify.global provides Bitcoin mining solutions and AI computational infrastructure, including data center, bare-metal, and cloud capabilities.

What you will do

  • Architect and maintain NVIDIA and AMD GPU device plugin and Kubernetes Operator integrations.
  • Configure and optimize RDMA, SR-IOV, RoCEv2, and InfiniBand networking for distributed AI training.
  • Build automated DCGM-based remediation pipelines to detect, isolate, and reset degraded GPU and NIC components.
  • Implement GPU slicing with MIG and vGPU for multi-tenant inference workloads.
  • Tune kernel parameters, device drivers, CUDA, and NCCL to optimize containerized AI workloads.
  • Collaborate with scheduling and storage teams on topology-aware placement, data movement, hardware standards, and complex performance investigations.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • 5+ years of systems engineering experience with strong proficiency in Linux kernel internals, C, or Go.
  • Hands-on experience with NVIDIA H100/A100 GPU architectures, CUDA runtimes, RDMA, and InfiniBand.
  • Deep understanding of containerized environments and Kubernetes device plugin architecture.
  • Experience operating, debugging, and scaling bare-metal systems in large-scale production or HPC environments.
  • Familiarity with Terraform, Ansible, and CI/CD infrastructure automation.

Nice to have

  • Experience in high-velocity, high-growth engineering environments.

Culture & Benefits

  • Full-time employment.
  • Collaborative work across infrastructure, scheduling, storage, and reliability engineering teams.
  • Opportunities to mentor team members and establish documentation standards for an evolving AI hardware stack.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →