5 дней назад
Senior Engineering Manager (AI/ML Inference)
250 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Engineering Manager (AI/ML Inference): Leading a team that makes large language models run faster, cheaper, and more reliably in production with an accent on inference optimization, serving architectures, and customer deployments. Focus on profiling vLLM, SGLang, and CUDA kernels, tuning latency, throughput, and cost, and moving workloads from proof of concept to monitored production services.
Location: San Francisco, CA, US; on-site
Compensation: $250,000–$300,000 per year plus bonus and Restricted Stock Units
Company
builds vertically integrated AI infrastructure across energy, data centers, cloud services, and machine learning workloads.
What you will do
- Lead and grow an engineering team focused on improving the speed, cost, and reliability of large language model inference in production.
- Design and optimize serving architectures, including prefill and decode disaggregation and request routing.
- Profile and improve the serving stack from vLLM and SGLang to CUDA kernels, analyzing latency, throughput, and cost bottlenecks.
- Adapt optimization methods across ML models and tailor deployments to customer models, traffic patterns, and operational constraints.
- Move workloads from proof of concept to monitored production services while partnering with customer engineering teams.
- Guide projects, technical strategy, performance goals, roadmaps, experiments, and production delivery while remaining hands-on with code and architecture.
Requirements
- At least 2 years of experience directly managing and leading an engineering team in a high-performance or ML-focused environment.
- Hands-on experience in software engineering, low-level optimization, or ML infrastructure, with production coding experience in Python or C++; Python is strongly preferred.
- Experience optimizing LLMs for high-throughput and low-latency inference and working with frameworks such as vLLM or SGLang.
- Experience profiling performance down to the kernel level and a strong understanding of GPU architecture and behavior.
- Working knowledge of AI/ML pipelines and the development and deployment lifecycle for ML models.
- Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field, plus strong customer and teammate communication skills.
Nice to have
- CUDA or comparable technology experience.
- Experience with Docker and Kubernetes.
- A record of making software systems faster and building or tuning AI/ML inference systems.
- Customer-facing experience with AI/ML projects.
Culture & Benefits
- Competitive compensation, bonus, equity, and Restricted Stock Units.
- Paid time off, holidays, leave programs, and parental leave.
- Health, dental, and vision insurance with employer HSA contributions.
- 401(k) retirement plan with company matching up to 4% of salary.
- Professional development, tuition reimbursement, wellness support, commuter benefits, meals allowance, and global travel insurance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Senior Machine Learning Engineer (AI)
250 000 - 350 000$
12 дней назад
Principal ML Engineer (AI)
172 000 - 258 000$
7 дней назад
Research Engineer (AI/ML)
195 000 - 400 000$
6 дней назад
Senior AI/ML Engineer
200 000 - 240 000$
10 дней назад
Senior AI Scientist
180 000 - 243 500$
7 дней назад
Engineering Manager (AI/ML)
208 000 - 260 000$