2 дня назад
Head of Performance Visibility (AI)
200 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Head of Performance Visibility (AI): Defining performance visibility for next-generation AI accelerator systems across custom silicon, compiler stacks, runtime libraries, and distributed inference environments with an accent on telemetry architecture, cross-layer event correlation, and performance modeling. Focus on building time-aligned analysis systems, identifying bottlenecks across chips, hosts, racks, and clusters, and transforming hardware signals into actionable insight for large-scale AI systems.
Location: On-site in San Jose, Santana Row; relocation support is available for candidates moving to San Jose.
Salary: $200,000–$300,000 plus significant equity.
Company
builds hardware for frontier intelligence by co-designing custom chips, racks, software, and manufacturing systems focused on high-throughput, low-latency AI inference.
What you will do
- Define the architecture for collecting and structuring telemetry across CPUs, drivers, interconnects, and multiple accelerators.
- Design scalable models for correlating performance events across device and host boundaries.
- Align hardware counters, runtime activity, communication phases, and workload semantics across multi-device systems.
- Develop counter taxonomies, derived performance models, and instrumentation strategies for future hardware generations.
- Build tools for identifying bottlenecks in multi-accelerator workloads and distributed inference across data center networks.
- Shape developer-facing analysis engines that turn raw telemetry into actionable insight for engineers debugging large-scale AI systems.
Requirements
- Deep experience building complex systems at the intersection of hardware and software.
- Hands-on experience building profiling, tracing, observability, telemetry, or performance analysis systems.
- Strong C++ or Rust programming skills with experience shipping low-level infrastructure close to hardware or runtime systems.
- Experience correlating time-series events across distributed systems, including timestamp synchronization and trace alignment.
- Experience with high-performance computing systems, large AI clusters, hardware counters, or performance modeling for ML workloads.
- Experience leading cross-functional architectural initiatives across hardware and software teams.
Culture & Benefits
- Fully in-person work in San Jose with collaboration across engineering and research disciplines.
- Medical, dental, and vision coverage with generous premium support, plus a $500 monthly credit for waiving medical benefits.
- $2,000 monthly housing subsidy for employees living within walking distance of the office.
- Wellness benefits, daily lunch and dinner, and an unlimited compute budget subject to ROI justification.
- Significant equity and relocation support for candidates moving to San Jose.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Principal Engineer (Kubernetes)
174 051 - 227 233$
2 дня назад
Principal Software Engineer (AI)
247 500 - 267 000$
2 дня назад
Principal Systems Software Engineer (AI Infrastructure)
260 000 - 340 000$
2 дня назад
Computer Architect (AI)
350 000 - 500 000$
CoreWeave
14 часов назад
Senior Software Engineer - AI Infrastructure Performance Insights & Observability
182 000 - 242 000$
2 дня назад
Enterprise Product Engineer (AI)
180 000 - 500 000$