Назад
Company hidden
обновлено 6 дней назад

AI Inference Engineer (AI)

Формат работы
remote (только Ukraine)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Ukraine
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Inference Engineer (AI/C++): Deploying and optimizing machine learning inference pipelines for edge devices with an accent on modern C++, AI model integration, memory usage, and latency. Focus on profiling production inference, integrating LLM and deep learning architectures, and transitioning models from research environments into existing products.

Location: Ukraine; remote flexibility is offered

Company

hirify.global is an AI engineering company building production systems for companies including Procter & Gamble and Shutterstock.

What you will do

  • Deploy machine learning models to edge devices using llama.cpp and ggml.
  • Integrate, evaluate, profile, and optimize AI inference pipelines for high-performance on-device execution.
  • Collaborate with researchers on coding, training, and transitioning models from research to production.
  • Integrate AI features into existing products using modern machine learning advancements.

Requirements

  • 4+ years of professional experience with Modern C++17/20.
  • Strong knowledge of memory management, multithreading, profiling, performance optimization, and low-level debugging.
  • Experience developing in Linux environments and integrating machine learning models into production applications.
  • Experience deploying and optimizing inference pipelines and profiling memory usage and latency.
  • Understanding of Transformer architecture, LLMs, diffusion models, tokenization, attention mechanisms, KV cache, quantization, model conversion, and deployment.
  • Practical experience with LLM deployment, computer vision, OCR, multimodal, speech, or image generation models.

Nice to have

  • Experience with llama.cpp, ggml, ONNX Runtime, TensorRT, TensorRT-LLM, OpenVINO, MLC LLM, ExecuTorch, or TVM.
  • CUDA, Vulkan Compute, Metal, OpenCL, TypeScript, or Python.
  • Experience contributing to open-source AI infrastructure projects.
  • Experience evaluating new models and integrating them into existing products.

Culture & Benefits

  • Remote work flexibility.
  • Competitive compensation with medical, wellness, and learning benefits.
  • English classes, professional development, well-being support, and career progression opportunities.
  • Ownership of problem-solving initiatives and support for responsible experimentation.
  • Regular meetups, tech talks, and collaborative support from colleagues.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →