4 дня назад
Staff GPU Inference SDET (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff GPU Inference SDET (AI): Building end-to-end release qualification and automated test systems for GPU inference stacks and multi-node accelerated compute fleets with an accent on distributed LLM serving, performance validation, numerical correctness, and fault resilience. Focus on designing workload benchmarks, validating prefill and decode optimization, simulating infrastructure failures, and integrating observability into continuous release gates.
Location: Hybrid in Sunnyvale, California, or Toronto, Canada
Company
builds large-scale AI computing hardware and infrastructure for high-speed model training and inference.
What you will do
- Design and scale automated release qualification systems, regression gates, and test pipelines for the GPU inference stack.
- Validate multi-node GPU cluster bring-up, model-serving workers, serving engines, container runtimes, drivers, and firmware.
- Benchmark distributed LLM serving workloads, including prefill and decode performance, continuous batching, prefix caching, KV-cache efficiency, and parallelism.
- Build workload replay and performance validation tools tracking TTFT, inter-token latency, throughput, tail latency, and capacity efficiency.
- Ensure numerical correctness, precision stability, determinism, and output quality across software and hardware changes.
- Develop fault-injection and chaos-testing suites for node failures, network degradation, memory leaks, driver or firmware mismatches, and recovery paths.
Requirements
- 8+ years of software engineering experience as an SDET, infrastructure quality lead, or systems test engineer.
- Hands-on experience provisioning and validating multi-node NVIDIA or AMD GPU clusters in public cloud or enterprise data center environments.
- Deep understanding of LLM serving engines, distributed runtimes, prefill and decode disaggregation, KV-cache management, and dynamic batching.
- Expert-level Python skills with experience building test automation frameworks, diagnostic tools, and CI/CD integrations.
- Strong proficiency with Kubernetes, Slurm, or Ray and high-performance interconnects such as InfiniBand, RoCE, or NCCL.
- Experience with root-cause analysis across software and hardware boundaries, stress testing, and distributed-system failure simulation.
Nice to have
- Experience with AMD ROCm/HIP or NVIDIA software stacks.
- Experience building workload replay tools, ML evaluation pipelines, or MLPerf Inference benchmark suites.
- Familiarity with PyTorch Profiler, NVTX, ROCm profilers, or C++.
Culture & Benefits
- Opportunity to build AI hardware and infrastructure beyond conventional GPU constraints.
- Access to cutting-edge AI research, open-source work, and high-performance AI supercomputing systems.
- Startup vitality combined with job stability.
- Non-corporate work culture focused on individual perspectives, learning, growth, and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Writer
9 дней назад
Software Quality Engineer (AI)
11 дней назад
Staff Software Engineer (AI)
140 400 - 372 300$
9 дней назад
Staff IT Engineer (AI)
200 000 - 240 000$
4 дня назад
Staff Software Engineer (AI Backend)
230 000 - 260 000$
Decagon
9 дней назад
Research Engineer (AI Safety)
200 000 - 400 000$
10 дней назад