Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Technical Support Engineer (AI): Diagnosing customer AI compute infrastructure issues across Linux hosts, GPUs, Kubernetes, Slurm, storage, and high-speed networking with an accent on hands-on troubleshooting and clear technical communication. Focus on investigating degraded training and inference workloads, building diagnostic tooling and runbooks, and resolving complex GPU cluster incidents.
Location: Las Vegas, Nevada, United States; on-site
Company
TensorWave provides a cloud platform for seamless, secure, reliable, and resilient AI compute at scale.
What you will do
- Own Level 2 tickets escalated from the global operations center through resolution or an evidence-based handoff to engineering.
- Diagnose Linux hosts, GPU health, Kubernetes workloads, Slurm scheduling, high-performance storage, and high-speed networking issues.
- Investigate degraded or failed AI training and inference workloads using logs, metrics, and cluster telemetry.
- Triage GPU and node hardware faults and coordinate remediation with data center operations.
- Communicate directly with customer engineering teams and document findings in runbooks and the knowledge base.
- Build diagnostic scripts and tools, partner with technical account managers, and participate in an on-call rotation.
Requirements
- 3+ years of experience in Linux systems administration, SRE, infrastructure operations, or technical support engineering with diagnostic ownership.
- Strong Linux troubleshooting skills covering systems, networking, file systems, storage, processes, resources, logs, and kernel-level issues.
- Working knowledge of Kubernetes cluster and pod state, scheduling, placement, and workload diagnosis.
- Experience with a batch scheduler in a shared compute environment, ideally Slurm.
- Scripting ability in Python or Bash and experience with structured incident or ticket-tracking systems such as JIRA or PagerDuty.
- Authorization to work in the United States is required. Willingness to participate in an on-call rotation is also required.
Nice to have
- Experience in GPU cloud, HPC, or AI/ML infrastructure operations.
- Familiarity with AMD Instinct, MI300X, MI325X, MI355X, ROCm, NVIDIA, or CUDA.
- Experience with RDMA, high-speed networking, Weka, Vast, container runtimes, Grafana, Prometheus, PyTorch, or JAX.
- Understanding of distributed training, inference-serving patterns, and ITIL or another structured incident management framework.
Culture & Benefits
- Stock options.
- 100% employee-paid medical, dental, and vision insurance.
- Health Savings Account contributions, Flexible Spending Account, and disability and life insurance options.
- 401(k), flexible PTO, paid holidays, and parental leave.
- Employee assistance, supplementary health benefits, and other in-office perks.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Customer Success Engineer
6 дней назад
Technical Support Engineer
125 000 - 200 000$
6 дней назад
Developer Support Engineer (AI)
75 000 - 135 000$
1 день назад
Technical Support Engineer (AI)
6 дней назад
Sr. Technical Support Engineer (AI)
6 дней назад
Technical Support Specialist (Robotics)
80 000 - 120 000$