1 день назад
Systems Operations Support Engineer (AI Infrastructure)
90 000 - 160 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Systems Operations Support Engineer (AI Infrastructure): Resolving escalated infrastructure issues across Linux, hardware, networking, Docker, NVIDIA CUDA, GPUs, and KVM virtualization with an accent on root-cause analysis and customer-facing technical support. Focus on diagnosing GPU and container failures, building Python and Bash automation, and creating runbooks that prevent recurring incidents.
Location: Westwood, Los Angeles office; fully on-site Monday–Friday, or four days on-site and one day working from home Sunday–Thursday
Salary: $90,000–$160,000 per year plus equity and benefits
Company
provides decentralized cloud computing infrastructure for AI projects and businesses.
What you will do
- Handle escalated support tickets involving GPU workloads, containers, networking, accounts, and host-side configuration.
- Support supplier onboarding and machine management through installation, configuration, and post-setup troubleshooting.
- Diagnose issues across Linux, Docker, NVIDIA CUDA and GPU drivers, KVM virtualization, and network infrastructure.
- Investigate performance problems involving GPU utilization, resource constraints, thermal throttling, driver conflicts, and disk I/O.
- Build Python and Bash diagnostic automation and maintain runbooks, escalation guides, and knowledge base articles.
- Collaborate with engineering and support teams to identify and document systemic platform issues.
Requirements
- Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions.
- Proficiency with Docker, Docker Compose, image management, cgroup limits, and Docker storage troubleshooting.
- Experience with Proxmox VE, VMware, or similar virtualization platforms, including VM provisioning and troubleshooting.
- Strong networking fundamentals covering VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting.
- Hands-on experience with NVIDIA GPU drivers, CUDA, GPU workload troubleshooting, Python, and Bash.
- Clear, professional written English and experience providing customer-facing or internal technical support.
Nice to have
- Experience with TensorFlow, PyTorch, and GPU-accelerated containers.
- Monitoring and observability experience with Prometheus or Grafana.
- RHCSA, CompTIA Linux+, or a similar certification.
- Experience using as a client or infrastructure supplier.
Culture & Benefits
- Early-stage startup environment focused on initiative, ownership, integrity, and continuous learning.
- Health, dental, vision, and life insurance.
- 401(k) with company match and meaningful equity.
- Onsite meals and snacks with close collaboration with founders and technical leaders.
Hiring process
- 15-minute virtual initial screening.
- 45-minute virtual experience interview.
- Two-hour onsite meet-and-greet and LLM-assisted Linux systems operations technical assessment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
IT Engineer
150 000 - 250 000$
1 день назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
1 день назад
Data Center Engineer (AI Infrastructure)
130 000 - 210 000$
1 день назад
Linux Device Management Engineer
160 000 - 200 000$
1 день назад
Infrastructure Engineer (AI)
6 дней назад
IT Operations Engineer (Cloud/Network)
115 000 - 155 000$