Назад
Company hidden
2 месяца назад

Machine Learning Engineer (LLM)

Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
Armenia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer (LLM) (AI inference infrastructure): Building and optimizing production systems for serving large language models with an accent on inference performance, model deployment, quantization, and GPU utilization. Focus on profiling workloads, designing scalable distributed inference architectures, and improving reliability, latency, throughput, and cost-efficiency.

Location: Yerevan, Armenia

Company

hirify.global develops production AI solutions and infrastructure for deploying and operating machine learning systems.

What you will do

  • Design, deploy, and optimize large language model inference systems for production.
  • Evaluate serving performance across latency, throughput, memory utilization, and cost-efficiency.
  • Implement model-serving solutions with vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or similar technologies.
  • Develop model conversion, deployment, benchmarking, and evaluation workflows.
  • Apply quantization techniques and profile GPU workloads to identify and resolve performance bottlenecks.
  • Collaborate with Platform, Infrastructure, and Product Engineering teams on scalable, reliable AI services and architecture decisions.

Requirements

  • 3+ years of ML engineering experience focused on model deployment, serving, and inference optimization.
  • Production experience with modern LLM serving frameworks such as vLLM, TensorRT-LLM, SGLang, or similar technologies.
  • Deep understanding of transformer architectures, attention mechanisms, KV caching, autoregressive generation, and prefill versus decode phases.
  • Experience with FP8, INT8, AWQ, GPTQ, or similar quantization techniques and their accuracy, latency, and hardware trade-offs.
  • Advanced Python skills, strong performance profiling experience, and working knowledge of PyTorch, ONNX, TorchScript, and TensorRT.
  • Ability to collaborate with platform, infrastructure, and product teams in production environments.

Nice to have

  • Knowledge of GPU architecture, HBM bandwidth, tensor cores, NVLink, and modern NVIDIA accelerators such as B200, H100, or L40S.
  • Experience with Kubernetes, containerized ML workloads, and Triton Inference Server.
  • Knowledge of distributed inference, tensor parallelism, NCCL, RDMA/InfiniBand, and GPUDirect.
  • Experience with speculative decoding, prefix caching, multi-tenant serving, or LoRA adapter hot-swapping.

Culture & Benefits

  • Full-time role focused on production AI infrastructure.
  • Cross-functional collaboration with platform, infrastructure, and product engineering teams.
  • Opportunity to evaluate emerging LLM serving, inference optimization, and AI infrastructure technologies.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →