1 день назад
Linux GPU Infrastructure Support Engineer
90 000 - 160 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Linux GPU Infrastructure Support Engineer (AI infrastructure/Linux/GPU): Troubleshooting complex Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, Docker, virtualization, networking, hardware, and GPU workloads with an accent on root-cause analysis and escalated technical support. Focus on building diagnostic tooling and runbooks, resolving recurring platform issues, and supporting TensorFlow and PyTorch workloads.
Location: Westwood, Los Angeles, United States. The role is primarily on-site, with an alternative Sunday–Thursday schedule including four on-site days and one work-from-home day.
Salary: $90,000–$160,000 per year, plus equity and benefits.
Company
provides decentralized cloud computing infrastructure for AI projects and businesses.
What you will do
- Diagnose and resolve issues across NVIDIA GPU drivers, CUDA, Docker, KVM virtualization, hardware, storage, and networking.
- Investigate GPU utilization, resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks, and workload failures.
- Handle complex escalated support tickets end-to-end, coordinating with clients, infrastructure suppliers, engineering, and host support teams.
- Support supplier onboarding and machine management, including installation, configuration, BIOS, firmware, drivers, and network setup.
- Build diagnostic and automation tooling in Python and Bash to reduce manual triage.
- Create runbooks, escalation guides, and knowledge base articles while identifying recurring platform issues.
Requirements
- Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, permissions, and command-line troubleshooting.
- Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and storage troubleshooting.
- Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors.
- Strong networking fundamentals covering VLANs, DNS, DHCP, NAT, VPNs, firewalls, and L2/L3 troubleshooting.
- Hands-on experience with NVIDIA GPU drivers, CUDA, GPU workloads, Python, and Bash scripting.
- Customer-facing or internal technical support experience, clear written English communication, and the ability to independently prioritize escalated tickets.
Nice to have
- Experience with TensorFlow, PyTorch, and GPU-accelerated containers.
- Monitoring and observability experience with Prometheus or Grafana.
- RHCSA, CompTIA Linux+, or a similar certification.
- Experience using the platform as a client or infrastructure supplier.
Culture & Benefits
- Comprehensive health, dental, vision, and life insurance.
- 401(k) with company match.
- Meaningful early-stage equity.
- On-site meals and snacks.
- Fast-paced startup environment with close collaboration with founders and technical leaders.
Hiring process
- 15-minute virtual initial screening.
- 45-minute virtual experience interview.
- Two-hour on-site meet-and-greet and LLM-assisted Linux systems operations technical assessment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
IT Engineer
150 000 - 250 000$
1 день назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
1 день назад
Linux Device Management Engineer
160 000 - 200 000$
3 дня назад
Systems Administrator, IT Infrastructure (Fintech)
80 000 - 110 000$
1 день назад
Data Center Engineer (AI Infrastructure)
130 000 - 210 000$
1 день назад