1 день назад
Technical Support Engineer (AI Infrastructure)
90 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Technical Support Engineer (AI Infrastructure): Diagnosing and resolving escalated infrastructure issues across Linux, Docker, NVIDIA CUDA/GPU systems, networking, and virtualization with an accent on root-cause analysis, customer support, and automation. Focus on building Python and Bash diagnostic tooling, troubleshooting GPU-accelerated workloads, and creating runbooks that eliminate recurring escalations.
Location: On-site in Westwood, Los Angeles, United States. Schedule: Sunday–Thursday, with participation in a defined on-call rotation that may include weekend coverage.
Salary: $90,000–$150,000 annually, plus equity and benefits.
Company
provides decentralized cloud computing infrastructure for AI projects and businesses.
What you will do
- Handle escalated support tickets involving GPU workload failures, containers, networking, account infrastructure, and host-side configuration.
- Diagnose issues across Ubuntu and other Linux environments, Docker, NVIDIA CUDA/GPU drivers, virtualization, storage, and network layers.
- Investigate GPU utilization, resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks, and connectivity failures.
- Support infrastructure suppliers with hardware installation, BIOS and firmware settings, driver configuration, networking, and machine management.
- Build Python and Bash diagnostic and automation tools to reduce manual triage and maintain runbooks, escalation guides, and knowledge base articles.
- Collaborate with engineering and infrastructure support teams on recurring platform issues, AI frameworks, and GPU-accelerated workloads.
Requirements
- Strong Linux SysOps experience with Ubuntu Server, RHEL/CentOS, or Debian, including systems, networking, storage, permissions, and command-line troubleshooting.
- Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage.
- Experience with Proxmox VE, VMware, or similar hypervisors and with provisioning and troubleshooting virtual machines.
- Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting.
- Experience with VLAN, DNS, DHCP, NAT, VPN, firewall rules, L2/L3 troubleshooting, Python, and Bash automation.
- Strong English written communication and experience providing customer-facing or internal technical support are required.
Nice to have
- Familiarity with TensorFlow, PyTorch, and GPU-accelerated containers.
- Monitoring and observability experience with Prometheus or Grafana.
- RHCSA, CompTIA Linux+, or a similar certification.
- Experience using the platform as a client or infrastructure supplier.
Culture & Benefits
- Health, dental, vision, and life insurance.
- 401(k) with company match.
- Meaningful early-stage equity.
- Onsite meals and snacks.
- Fast-paced startup environment with close collaboration with founders and technical leaders.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
IT Engineer
150 000 - 250 000$
1 день назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
1 день назад
Data Center Engineer (AI Infrastructure)
130 000 - 210 000$
1 день назад
Linux Device Management Engineer
160 000 - 200 000$
1 день назад
Infrastructure Engineer (AI)
1 день назад