обновлено 7 дней назад
Senior Solutions Engineer (AI/GPU Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Solutions Engineer (AI/GPU Infrastructure): Resolving complex production escalations for large-scale AI compute customers with an accent on Kubernetes, GPU workloads, kernel-level debugging, and high-performance networking. Focus on diagnosing P1 failures, developing remediation tools, translating customer issues into product improvements, and strengthening platform reliability.
Location: Remote, Las Vegas, Nevada, United States. Authorization to work in the United States is required.
Company
TensorWave provides a secure, reliable, and resilient cloud platform for AI compute at scale.
What you will do
- Resolve complex technical escalations beyond Global Operations Center runbooks through code-level debugging and architectural investigation.
- Partner with customer technical leads to diagnose production issues and coordinate rapid resolutions.
- Develop diagnostic scripts, remediation tools, and workarounds while permanent patches are being built.
- Own end-to-end P1 incident resolution and deliver actionable post-incident analysis with Technical Account Managers.
- Translate recurring customer issues into evidence-based feature requests and product improvements.
- Document platform behavior and improve operational runbooks.
Requirements
- 5–9 years of experience in infrastructure engineering, platform engineering, or SRE, focused on high-performance computing or large-scale AI stacks.
- Deep Kubernetes administration experience, including scheduler internals and controller code.
- Experience orchestrating GPU workloads and diagnosing training failures with ROCm or CUDA.
- Strong knowledge of RDMA/RoCEv2, SR-IOV, BGP, switch telemetry, Linux kernel networking, hugepages, and cgroups.
- Proficiency in Python and Ansible for building diagnostic and remediation tools.
- Strong technical communication skills when presenting findings to engineering leadership.
Nice to have
- Customer-facing engineering experience in Solutions Engineering or Technical Support Engineering.
- Experience supporting high-uptime environments requiring 24/7/365 availability.
Culture & Benefits
- Stock options.
- 100% paid medical, dental, and vision insurance for employees.
- Health Savings Account contributions, Flexible Spending Account, and supplemental insurance options.
- 401(k), flexible PTO, paid holidays, parental leave, and disability insurance.
- Employee Assistance Program and supplementary health benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Storage Engineer (AI Infrastructure)
150 000 - 300 000$
1 день назад
Senior DevOps Engineer (Cloud)
155 000 - 175 000$
8 дней назад
Senior DevSecOps Engineer (AWS)
120 000 - 150 000$
7 дней назад
DevOps Engineer (AI)
90 000 - 130 000$
Writer
8 дней назад
Infrastructure Engineer (AI)
9 дней назад