10 часов назад
LLM Inference Deployment Engineer (AI)
180 000 - 240 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
LLM Inference Deployment Engineer (AI): Deploying and optimizing large language models for high-performance inference on energy-efficient AI accelerators with an accent on runtime execution, model integration, and low-latency serving. Focus on building containerized inference pipelines, optimizing batching, caching, tensor parallelism, and memory usage for real-time LLM applications.
Location: Remote in the United States or Canada
Salary: $180,000–$240,000 USD per year; $175,000–$245,000 CAD per year
Company
develops advanced AI hardware and software systems for efficient edge-to-cloud computing, using in-memory computing technology for power-, energy-, and space-constrained applications.
What you will do
- Deploy and optimize post-trained large language models from libraries such as Hugging Face.
- Use inference runtimes including ONNX Runtime and vLLM to improve execution efficiency.
- Optimize batching, caching, and tensor parallelism for scalable real-time inference.
- Develop and maintain high-performance inference pipelines with Docker, Kubernetes, and inference servers.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- Experience with LLM inference deployment, model optimization, and runtime engineering.
- Strong expertise in PyTorch, ONNX Runtime, vLLM, TensorRT-LLM, and DeepSpeed.
- Advanced Python skills for model integration and performance tuning.
- Knowledge of model representations, framework-level optimization, and LLM memory optimization for long-context applications.
- Experience with Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, or TorchServe, as well as real-time LLM applications such as chatbots, code generation, and retrieval-augmented generation.
Culture & Benefits
- Work remotely from the United States or Canada.
- Contribute to AI hardware and software systems spanning edge-to-cloud computing.
- Join a company launched in 2022 and led by technologists with semiconductor design and AI systems experience.
- Equal employment opportunity employer in the United States.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 часов назад
Inference Runtime Engineer (AI)
200 000 - 400 000$
4 дня назад
Senior AI Systems Engineer (Inference)
6 часов назад
Developer Relations Engineer (AI Inference)
200 000 - 400 000$
6 дней назад
Lead AI Scientist (LLM)
220 000 - 260 000$
1 день назад
Forward Deployed Engineer (Post-Sales) (AI)
230 000 - 300 000$
7 часов назад
Training Infrastructure Engineer (AI)
210 000 - 320 000$