Назад
Company hidden
1 день назад

Technical Support Engineer (AI Infrastructure)

90 000 - 150 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Support Engineer (AI Infrastructure): Diagnosing and resolving escalated infrastructure issues across Linux, Docker, NVIDIA CUDA/GPU systems, networking, and virtualization with an accent on root-cause analysis, customer support, and automation. Focus on building Python and Bash diagnostic tooling, troubleshooting GPU-accelerated workloads, and creating runbooks that eliminate recurring escalations.

Location: On-site in Westwood, Los Angeles, United States. Schedule: Sunday–Thursday, with participation in a defined on-call rotation that may include weekend coverage.

Salary: $90,000–$150,000 annually, plus equity and benefits.

Company

hirify.global provides decentralized cloud computing infrastructure for AI projects and businesses.

What you will do

  • Handle escalated support tickets involving GPU workload failures, containers, networking, account infrastructure, and host-side configuration.
  • Diagnose issues across Ubuntu and other Linux environments, Docker, NVIDIA CUDA/GPU drivers, virtualization, storage, and network layers.
  • Investigate GPU utilization, resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks, and connectivity failures.
  • Support infrastructure suppliers with hardware installation, BIOS and firmware settings, driver configuration, networking, and machine management.
  • Build Python and Bash diagnostic and automation tools to reduce manual triage and maintain runbooks, escalation guides, and knowledge base articles.
  • Collaborate with engineering and infrastructure support teams on recurring platform issues, AI frameworks, and GPU-accelerated workloads.

Requirements

  • Strong Linux SysOps experience with Ubuntu Server, RHEL/CentOS, or Debian, including systems, networking, storage, permissions, and command-line troubleshooting.
  • Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage.
  • Experience with Proxmox VE, VMware, or similar hypervisors and with provisioning and troubleshooting virtual machines.
  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting.
  • Experience with VLAN, DNS, DHCP, NAT, VPN, firewall rules, L2/L3 troubleshooting, Python, and Bash automation.
  • Strong English written communication and experience providing customer-facing or internal technical support are required.

Nice to have

  • Familiarity with TensorFlow, PyTorch, and GPU-accelerated containers.
  • Monitoring and observability experience with Prometheus or Grafana.
  • RHCSA, CompTIA Linux+, or a similar certification.
  • Experience using the hirify.global platform as a client or infrastructure supplier.

Culture & Benefits

  • Health, dental, vision, and life insurance.
  • 401(k) with company match.
  • Meaningful early-stage equity.
  • Onsite meals and snacks.
  • Fast-paced startup environment with close collaboration with founders and technical leaders.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →