Назад
Company hidden
1 день назад

Systems Operations Support Engineer (AI Infrastructure)

90 000 - 160 000$
Формат работы
onsite/hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Systems Operations Support Engineer (AI Infrastructure): Resolving escalated infrastructure issues across Linux, hardware, networking, Docker, NVIDIA CUDA, GPUs, and KVM virtualization with an accent on root-cause analysis and customer-facing technical support. Focus on diagnosing GPU and container failures, building Python and Bash automation, and creating runbooks that prevent recurring incidents.

Location: Westwood, Los Angeles office; fully on-site Monday–Friday, or four days on-site and one day working from home Sunday–Thursday

Salary: $90,000–$160,000 per year plus equity and benefits

Company

hirify.global provides decentralized cloud computing infrastructure for AI projects and businesses.

What you will do

  • Handle escalated support tickets involving GPU workloads, containers, networking, accounts, and host-side configuration.
  • Support supplier onboarding and machine management through installation, configuration, and post-setup troubleshooting.
  • Diagnose issues across Linux, Docker, NVIDIA CUDA and GPU drivers, KVM virtualization, and network infrastructure.
  • Investigate performance problems involving GPU utilization, resource constraints, thermal throttling, driver conflicts, and disk I/O.
  • Build Python and Bash diagnostic automation and maintain runbooks, escalation guides, and knowledge base articles.
  • Collaborate with engineering and support teams to identify and document systemic platform issues.

Requirements

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions.
  • Proficiency with Docker, Docker Compose, image management, cgroup limits, and Docker storage troubleshooting.
  • Experience with Proxmox VE, VMware, or similar virtualization platforms, including VM provisioning and troubleshooting.
  • Strong networking fundamentals covering VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting.
  • Hands-on experience with NVIDIA GPU drivers, CUDA, GPU workload troubleshooting, Python, and Bash.
  • Clear, professional written English and experience providing customer-facing or internal technical support.

Nice to have

  • Experience with TensorFlow, PyTorch, and GPU-accelerated containers.
  • Monitoring and observability experience with Prometheus or Grafana.
  • RHCSA, CompTIA Linux+, or a similar certification.
  • Experience using hirify.global as a client or infrastructure supplier.

Culture & Benefits

  • Early-stage startup environment focused on initiative, ownership, integrity, and continuous learning.
  • Health, dental, vision, and life insurance.
  • 401(k) with company match and meaningful equity.
  • Onsite meals and snacks with close collaboration with founders and technical leaders.

Hiring process

  • 15-minute virtual initial screening.
  • 45-minute virtual experience interview.
  • Two-hour onsite meet-and-greet and LLM-assisted Linux systems operations technical assessment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →