1 день назад
Forward Deployed Engineer (Inference)
180 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Forward Deployed Engineer (Inference) (AI Infrastructure/LLM Inference): Building and operating production inference solutions for customer workloads with an accent on profiling, benchmarking, tracing, and systems engineering. Focus on optimizing latency, throughput, reliability, and cost across inference engines, model-serving infrastructure, accelerators, and observability tooling.
Location: San Francisco, CA; on-site
Salary: $180,000–$300,000 per year
Company
AI infrastructure company building systems that optimize inference workloads across different types of silicon for high-volume, mission-critical AI applications.
What you will do
- Own customer engagements from technical handoff through evaluation, deployment, production operation, and expansion.
- Profile, benchmark, and trace customer workloads using real models and traffic.
- Design and lead technical evaluations, explaining methodologies, results, constraints, and tradeoffs to engineering teams.
- Implement inference workloads in production across inference engines, model serving, infrastructure, observability, and performance tooling.
- Own production performance and reliability, including on-call support and direct debugging of latency, throughput, reliability, and error-rate regressions.
- Maintain technical relationships with customer engineering teams and turn recurring problems into product improvements.
Requirements
- Strong systems-engineering fundamentals and experience owning technical problems through production deployment.
- Understanding of LLM inference fundamentals, including prefill, decode, latency, throughput, and performance tradeoffs.
- Experience profiling, benchmarking, and tracing complex systems, reading traces, and identifying performance bottlenecks.
- Ability to work across unfamiliar systems, operate under ambiguity and urgency, and turn incomplete requirements into working implementations.
- Strong technical communication skills with highly technical customers and engineering audiences.
- Strong production ownership and debugging instincts.
Nice to have
- Experience with LLM inference engines, model-serving infrastructure, GPU or accelerator optimization, or distributed inference.
- Experience with CUDA, Triton, TensorRT-LLM, vLLM, SGLang, Kubernetes, GPU profiling, or tracing tools.
- Customer-facing engineering, solutions architecture, forward-deployed engineering, or latency-sensitive production experience.
Culture & Benefits
- Highly technical, hands-on environment with significant individual ownership and autonomy.
- Broad problem-solving across the stack rather than narrowly defined coding responsibilities.
- Fast-moving production environment focused on benchmarking, profiling, debugging, implementation, and operations.
- Team values include resourcefulness, technical ability, high standards, collaboration, emotional intelligence, rapid learning, and first-principles thinking.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Anthropic
3 дня назад
Performance Engineer (AI)
280 000 - 850 000$
7 дней назад
Senior Engineering Manager (AI/ML Inference)
250 000 - 300 000$
7 дней назад
AI Performance Engineer
100 000 - 150 000$
6 дней назад
Inference Performance Engineer (AI)
Microsoft AI
1 день назад
Member of Technical Staff, Inference Systems Research (AI)
142 800 - 274 800$
Baseten
3 дня назад
Forward Deployed Engineers (AI)
200 000 - 400 000$