13 часов назад
Senior Performance Engineer (AI)
135 000 - 170 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Performance Engineer (AI): Building roofline models, benchmarks, and automated test infrastructure to measure scale-up fabric performance on GPU inference and training workloads with an accent on bottleneck analysis, workload scalability, and competitive benchmarking. Focus on analyzing 16-to-32+ GPU clusters, debugging performance across hardware and software boundaries, and translating results into architecture, firmware, product, and customer-facing decisions.
Location: San Jose, California, United States; on-site
Salary: $135,000–$170,000 per year, depending on experience, level, and business need; discretionary bonus, incentives, and benefits may apply.
Company
provides rack-scale AI infrastructure through connectivity solutions integrating CXL, Ethernet, NVLink, PCIe, UALink, and software technologies.
What you will do
- Establish theoretical and measured roofline models and comparative benchmarks for scale-up fabric performance.
- Build GPU-cluster benchmarks with tools including NVBandwidth and NCCL across configurations and switch topologies.
- Run inference workloads and evaluate performance scaling from 16 to 32 GPUs and beyond.
- Analyze bottlenecks, compare competing fabric switch solutions, and develop data-driven performance differentiation.
- Design automated lab infrastructure, test pipelines, traffic-generation tools, and reporting systems.
- Partner with ASIC architecture, firmware, software, product, applications, and marketing teams to influence decisions and support customer engagements.
Requirements
- Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field.
- Recent graduates with directly relevant projects, research, or internships are considered; alternatively, 2–5 years of industry experience in performance or systems engineering.
- Hands-on experience running and analyzing AI/ML workloads on GPU or accelerator clusters.
- Ability to debug and root-cause system-level performance issues across hardware, firmware, software, and network boundaries.
- Knowledge of computer systems, GPU systems, datacenter networking, PCIe, Ethernet, CUDA, MPI, collective communication libraries, drivers, and OS-level performance tools.
- Proficiency in Python or similar scripting and automation tools for test pipelines and performance-data analysis.
Nice to have
- MS or PhD in a related technical field.
- Experience with scale-up fabrics and UALink, PCIe Gen 6/Gen 7, Ethernet, or UEC.
- Understanding of LLM, MoE, recommender-system, inference, and training workload communication patterns.
- Experience developing roofline models and competitive analyses for switching, networking, or accelerator silicon.
- Strong technical writing and communication skills for executive, customer, and marketing audiences.
Culture & Benefits
- Work on rack-scale AI infrastructure and next-generation connectivity technologies.
- Collaborate across architecture, firmware, software, product, applications, and marketing functions.
- Potential eligibility for discretionary bonus, incentives, and benefits.
- Applications are encouraged from candidates with diverse backgrounds and experiences.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 часов назад
GPU Systems Engineer (AI)
150 000 - 300 000$
9 часов назад
Senior Storage Infrastructure Engineer (AI)
100 000 - 300 000$
2 дня назад
Network Operations Engineer (AI)
173 000 - 279 000$
3 дня назад
Senior IT Operations Engineer (AI)
150 000 - 190 000$
11 часов назад
Senior Network Engineer (AI/AIOps)
105 000 - 125 000$
6 дней назад
Senior Network Engineer (InfiniBand / UFM)
170 000 - 210 000$