4 дня назад
AI Inference Engineer (LLM)
215 000 - 260 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Inference Engineer (LLM) (AI/Cloud Infrastructure): Building and optimizing production inference systems for large language models with an accent on serving architecture, GPU performance, latency, throughput, and cost efficiency. Focus on profiling vLLM and SGLang down to CUDA kernels, adapting deployments to customer workloads, and shipping reliable monitored services from proof of concept to production.
Location: San Francisco, CA, United States; on-site
Salary: $215,000–$260,000 per year plus bonus; Restricted Stock Units included
Company
builds vertically integrated energy and AI infrastructure, operating across energy, data center construction, cloud services, and machine learning workloads.
What you will do
- Bring modern large language model inference techniques into production and refine them for real customer workloads.
- Design and optimize serving architectures, including prefill and decode disaggregation and request routing.
- Profile and improve the serving stack from vLLM and SGLang down to CUDA kernels.
- Tune deployments for latency, throughput, cost, reliability, and real-world traffic.
- Partner with customer engineering teams to move workloads from proof of concept to monitored production services.
- Build, specify, test, and ship inference software and product features end to end.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Production software development experience with Python or C++; Python is strongly preferred.
- Hands-on experience optimizing large language models for high-throughput and low-latency inference.
- Experience with vLLM or SGLang, performance profiling, kernel-level analysis, and GPU architecture.
- Knowledge of AI/ML pipelines and the full process of developing and deploying machine learning models.
- Strong communication skills for explaining complex technical topics to customers and engineering teammates.
Nice to have
- Experience with CUDA or comparable technologies.
- Experience building or tuning AI/ML inference systems.
- Experience with Docker and Kubernetes.
- Customer-facing AI/ML project experience.
Culture & Benefits
- Competitive compensation, equity, and Restricted Stock Units.
- Paid time off, holidays, leave programs, and parental leave.
- Health, dental, vision, HSA, life, and disability insurance.
- Professional development, tuition reimbursement, and mental health support.
- 401(k) plan with company matching up to 4% of salary.
- Commuter benefits, meals allowance, cell phone stipend, volunteer time off, and global travel insurance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Software Engineer (AI/ML Tools)
97 700 - 144 400$
11 дней назад
Senior ML Infrastructure Engineer (Autonomous Driving)
128 700 - 261 300$
3 дня назад
AI Platform Engineer
130 000 - 180 000$
10 дней назад
Machine Learning Engineer, Infra, AI for Drug Discovery
147 600 - 274 000$
11 дней назад
Distinguished Engineer - AI (AI Infrastructure)
266 050 - 396 000$
9 дней назад
AI/ML Ops and Data Engineer (GCP)
134 800 - 195 000$