5 часов назад
Distributed LLM Inference Engineer (AI)
170 000 - 245 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Distributed LLM Inference Engineer (AI): Building and optimizing batch and online inference solutions for large-scale machine learning with an accent on distributed systems, Ray Data integration, and open-source LLM engines. Focus on achieving low-cost, high-throughput, low-latency inference, integrating vLLM and contributing improvements to deep learning infrastructure.
Location: Hybrid in San Francisco or Palo Alto, United States
Salary: $170,000–$245,000 annual target base salary, plus equity
Company
commercializes Ray, an open-source distributed computing platform for scalable machine learning applications.
What you will do
- Develop and ship end-to-end batch and online inference solutions for open-source Ray users and customers.
- Integrate Ray Data with LLM engines and implement optimizations for large-scale machine learning inference.
- Integrate open-source technologies such as vLLM into solutions and contribute improvements upstream.
- Implement and extend best practices from the open-source and research communities.
Requirements
- Experience running machine learning inference at large scale with high throughput and low latency.
- Familiarity with deep learning and frameworks such as PyTorch.
- Solid understanding of distributed systems and machine learning inference challenges.
- Ability to work in a hybrid arrangement in San Francisco or Palo Alto.
Nice to have
- Machine learning systems knowledge and experience using Ray.
- Experience with LLM engines such as vLLM or TensorRT-LLM.
- Contributions to PyTorch, TensorFlow, Triton, TVM, or MLIR.
- Prior experience working with GPUs or CUDA.
Culture & Benefits
- Stock options and equity participation.
- Healthcare premiums covered at 99% for employees and dependents.
- 401(k) retirement plan, education and wellbeing stipend, and paid parental leave.
- Fertility benefits, paid time off, commute reimbursement, and covered in-office meals.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
Inference Engineer (AI)
195 000 - 285 000$
6 часов назад
Member of Technical Staff - ML Systems & Inference
250 000 - 350 000$
5 часов назад
ML Performance Engineer
200 000 - 350 000$
5 часов назад
Software Engineer (AI Inference)
175 000 - 275 000$
5 дней назад
AI Infrastructure Engineer (GPU)
170 500 - 315 490$
4 часа назад
Principal System Software Engineer, AI Inference Execution
195 000 - 285 000$