Назад
обновлено 7 дней назад

Senior Solutions Engineer (AI/GPU Infrastructure)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Solutions Engineer (AI/GPU Infrastructure): Resolving complex production escalations for large-scale AI compute customers with an accent on Kubernetes, GPU workloads, kernel-level debugging, and high-performance networking. Focus on diagnosing P1 failures, developing remediation tools, translating customer issues into product improvements, and strengthening platform reliability.

Location: Remote, Las Vegas, Nevada, United States. Authorization to work in the United States is required.

Company

TensorWave provides a secure, reliable, and resilient cloud platform for AI compute at scale.

What you will do

  • Resolve complex technical escalations beyond Global Operations Center runbooks through code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues and coordinate rapid resolutions.
  • Develop diagnostic scripts, remediation tools, and workarounds while permanent patches are being built.
  • Own end-to-end P1 incident resolution and deliver actionable post-incident analysis with Technical Account Managers.
  • Translate recurring customer issues into evidence-based feature requests and product improvements.
  • Document platform behavior and improve operational runbooks.

Requirements

  • 5–9 years of experience in infrastructure engineering, platform engineering, or SRE, focused on high-performance computing or large-scale AI stacks.
  • Deep Kubernetes administration experience, including scheduler internals and controller code.
  • Experience orchestrating GPU workloads and diagnosing training failures with ROCm or CUDA.
  • Strong knowledge of RDMA/RoCEv2, SR-IOV, BGP, switch telemetry, Linux kernel networking, hugepages, and cgroups.
  • Proficiency in Python and Ansible for building diagnostic and remediation tools.
  • Strong technical communication skills when presenting findings to engineering leadership.

Nice to have

  • Customer-facing engineering experience in Solutions Engineering or Technical Support Engineering.
  • Experience supporting high-uptime environments requiring 24/7/365 availability.

Culture & Benefits

  • Stock options.
  • 100% paid medical, dental, and vision insurance for employees.
  • Health Savings Account contributions, Flexible Spending Account, and supplemental insurance options.
  • 401(k), flexible PTO, paid holidays, parental leave, and disability insurance.
  • Employee Assistance Program and supplementary health benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →