Назад
Company hidden
10 часов назад

LLM Inference Deployment Engineer (AI)

180 000 - 240 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
LLM Inference Deployment Engineer (AI): Deploying and optimizing large language models for high-performance inference on energy-efficient AI accelerators with an accent on runtime execution, model integration, and low-latency serving. Focus on building containerized inference pipelines, optimizing batching, caching, tensor parallelism, and memory usage for real-time LLM applications.

Location: Remote in the United States or Canada

Salary: $180,000–$240,000 USD per year; $175,000–$245,000 CAD per year

Company

hirify.global develops advanced AI hardware and software systems for efficient edge-to-cloud computing, using in-memory computing technology for power-, energy-, and space-constrained applications.

What you will do

  • Deploy and optimize post-trained large language models from libraries such as Hugging Face.
  • Use inference runtimes including ONNX Runtime and vLLM to improve execution efficiency.
  • Optimize batching, caching, and tensor parallelism for scalable real-time inference.
  • Develop and maintain high-performance inference pipelines with Docker, Kubernetes, and inference servers.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • Experience with LLM inference deployment, model optimization, and runtime engineering.
  • Strong expertise in PyTorch, ONNX Runtime, vLLM, TensorRT-LLM, and DeepSpeed.
  • Advanced Python skills for model integration and performance tuning.
  • Knowledge of model representations, framework-level optimization, and LLM memory optimization for long-context applications.
  • Experience with Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, or TorchServe, as well as real-time LLM applications such as chatbots, code generation, and retrieval-augmented generation.

Culture & Benefits

  • Work remotely from the United States or Canada.
  • Contribute to AI hardware and software systems spanning edge-to-cloud computing.
  • Join a company launched in 2022 and led by technologists with semiconductor design and AI systems experience.
  • Equal employment opportunity employer in the United States.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →