Назад
19 часов назад

Technical Support Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Support Engineer (AI): Diagnosing customer AI compute infrastructure issues across Linux hosts, GPUs, Kubernetes, Slurm, storage, and high-speed networking with an accent on hands-on troubleshooting and clear technical communication. Focus on investigating degraded training and inference workloads, building diagnostic tooling and runbooks, and resolving complex GPU cluster incidents.

Location: Las Vegas, Nevada, United States; on-site

Company

TensorWave provides a cloud platform for seamless, secure, reliable, and resilient AI compute at scale.

What you will do

  • Own Level 2 tickets escalated from the global operations center through resolution or an evidence-based handoff to engineering.
  • Diagnose Linux hosts, GPU health, Kubernetes workloads, Slurm scheduling, high-performance storage, and high-speed networking issues.
  • Investigate degraded or failed AI training and inference workloads using logs, metrics, and cluster telemetry.
  • Triage GPU and node hardware faults and coordinate remediation with data center operations.
  • Communicate directly with customer engineering teams and document findings in runbooks and the knowledge base.
  • Build diagnostic scripts and tools, partner with technical account managers, and participate in an on-call rotation.

Requirements

  • 3+ years of experience in Linux systems administration, SRE, infrastructure operations, or technical support engineering with diagnostic ownership.
  • Strong Linux troubleshooting skills covering systems, networking, file systems, storage, processes, resources, logs, and kernel-level issues.
  • Working knowledge of Kubernetes cluster and pod state, scheduling, placement, and workload diagnosis.
  • Experience with a batch scheduler in a shared compute environment, ideally Slurm.
  • Scripting ability in Python or Bash and experience with structured incident or ticket-tracking systems such as JIRA or PagerDuty.
  • Authorization to work in the United States is required. Willingness to participate in an on-call rotation is also required.

Nice to have

  • Experience in GPU cloud, HPC, or AI/ML infrastructure operations.
  • Familiarity with AMD Instinct, MI300X, MI325X, MI355X, ROCm, NVIDIA, or CUDA.
  • Experience with RDMA, high-speed networking, Weka, Vast, container runtimes, Grafana, Prometheus, PyTorch, or JAX.
  • Understanding of distributed training, inference-serving patterns, and ITIL or another structured incident management framework.

Culture & Benefits

  • Stock options.
  • 100% employee-paid medical, dental, and vision insurance.
  • Health Savings Account contributions, Flexible Spending Account, and disability and life insurance options.
  • 401(k), flexible PTO, paid holidays, and parental leave.
  • Employee assistance, supplementary health benefits, and other in-office perks.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →