Назад
Company hidden
1 день назад

Linux GPU Infrastructure Support Engineer

90 000 - 160 000$
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Linux GPU Infrastructure Support Engineer (AI infrastructure/Linux/GPU): Troubleshooting complex Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, Docker, virtualization, networking, hardware, and GPU workloads with an accent on root-cause analysis and escalated technical support. Focus on building diagnostic tooling and runbooks, resolving recurring platform issues, and supporting TensorFlow and PyTorch workloads.

Location: Westwood, Los Angeles, United States. The role is primarily on-site, with an alternative Sunday–Thursday schedule including four on-site days and one work-from-home day.

Salary: $90,000–$160,000 per year, plus equity and benefits.

Company

hirify.global provides decentralized cloud computing infrastructure for AI projects and businesses.

What you will do

  • Diagnose and resolve issues across NVIDIA GPU drivers, CUDA, Docker, KVM virtualization, hardware, storage, and networking.
  • Investigate GPU utilization, resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks, and workload failures.
  • Handle complex escalated support tickets end-to-end, coordinating with clients, infrastructure suppliers, engineering, and host support teams.
  • Support supplier onboarding and machine management, including installation, configuration, BIOS, firmware, drivers, and network setup.
  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage.
  • Create runbooks, escalation guides, and knowledge base articles while identifying recurring platform issues.

Requirements

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, permissions, and command-line troubleshooting.
  • Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and storage troubleshooting.
  • Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors.
  • Strong networking fundamentals covering VLANs, DNS, DHCP, NAT, VPNs, firewalls, and L2/L3 troubleshooting.
  • Hands-on experience with NVIDIA GPU drivers, CUDA, GPU workloads, Python, and Bash scripting.
  • Customer-facing or internal technical support experience, clear written English communication, and the ability to independently prioritize escalated tickets.

Nice to have

  • Experience with TensorFlow, PyTorch, and GPU-accelerated containers.
  • Monitoring and observability experience with Prometheus or Grafana.
  • RHCSA, CompTIA Linux+, or a similar certification.
  • Experience using the hirify.global platform as a client or infrastructure supplier.

Culture & Benefits

  • Comprehensive health, dental, vision, and life insurance.
  • 401(k) with company match.
  • Meaningful early-stage equity.
  • On-site meals and snacks.
  • Fast-paced startup environment with close collaboration with founders and technical leaders.

Hiring process

  • 15-minute virtual initial screening.
  • 45-minute virtual experience interview.
  • Two-hour on-site meet-and-greet and LLM-assisted Linux systems operations technical assessment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →