14 часов назад
Inference Systems Performance Architect (AI)
245 000 - 325 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Inference Systems Performance Architect (AI): Building workload-capture, benchmarking, performance-modeling, and simulation capabilities for large-scale LLM inference with an accent on heterogeneous, disaggregated inference, profiling, and customer SLOs. Focus on modeling future hardware configurations, localizing bottlenecks across hosts, accelerators, and fabric, and solving ambiguous performance challenges across engineering teams.
Location: San Jose, California, United States
Salary: $245,000–$325,000 USD base salary per year
Company
develops a full-stack generative AI platform for enterprise and government organizations, spanning AI chips, models, and cloud or on-premises deployment.
What you will do
- Define the technical strategy for end-to-end performance of large-scale LLM inference.
- Build reproducible workload-capture and agentic-benchmarking capabilities based on representative production traffic.
- Develop performance models and simulations to support capacity planning, customer SLOs, and next-generation hardware planning.
- Drive profiling tools that localize bottlenecks across hosts, accelerators, storage, networking, and the fabric.
- Guide trade-offs across reliability, scalability, operational cost, and ease of adoption while influencing model-optimization, systems, hardware, and product teams.
- Mentor senior and principal engineers, represent ’s performance capabilities to customers and partners, and resolve novel cross-functional challenges.
Requirements
- 12+ years of experience in performance engineering with technical leadership on large-scale, complex systems.
- Deep expertise in end-to-end performance analysis of distributed systems and bottleneck localization.
- Proven experience with realistic workload generation, simulation, and performance modeling calibrated against variable real-world workloads.
- Ability to apply performance engineering methods in unfamiliar domains and lead cross-functional initiatives.
- Experience mentoring senior engineers, influencing organizational direction, and representing an organization to customers and partners.
- Track record of independently delivering high-complexity, high-ambiguity work with significant product or roadmap impact.
Nice to have
- Experience with LLM inference serving, including continuous batching, prompt or KV caching, prefill/decode disaggregation, and tail-latency SLOs.
- Familiarity with inference simulation frameworks or agentic benchmarking.
- Public technical contributions through talks, writing, or systems-performance community participation.
Culture & Benefits
- Full-time US employment with equity and a competitive total rewards package.
- Medical insurance with 95% employee premium coverage and 77% dependent premium coverage.
- Health Savings Account with employer contribution, Flexible Spending Account options, dental, vision, disability, life, and AD&D insurance.
- Well-being benefits including Headspace, Gympass+, One Medical, and counseling through an Employee Assistance Program.
- Equal opportunity employment practices.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Head of Performance Visibility (AI)
200 000 - 300 000$
5 дней назад
Principal AI Systems Architect
195 000 - 285 000$
5 дней назад
Distinguished Software Engineer, Data Platform
300 000 - 320 000$
5 дней назад
Principal Architect - Performance Analysis and Modeling
195 000 - 285 000$
5 дней назад
Principal Systems Software Engineer (AI Infrastructure)
260 000 - 340 000$
5 дней назад
Principal Architect, Simulation Platform (AI)
200 000 - 240 000$