Назад
Company hidden
5 дней назад

Software Python Engineer (Inference)

Формат работы
remote (только Europe)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Serbia/Poland/Cyprus
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Python Engineer (Inference) (AI/Python): Building and improving the inference layer of the Gcore Inference platform with an accent on model deployment, framework integration, and GPU efficiency. Focus on debugging performance across software, infrastructure, hardware, and Kubernetes, while optimizing latency, throughput, memory use, utilization, and cost.

Location: Poland, Serbia, or Cyprus; hybrid or remote options may be available depending on the role. Work from anywhere in the world is available for up to 45 days per year.

Company

hirify.global provides infrastructure and software solutions for AI, cloud, network, and security across edge locations, cloud regions, and GPU platforms.

What you will do

  • Build and improve the inference layer of the hirify.global Inference platform.
  • Integrate and operate inference frameworks including vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM.
  • Bring new language and multimodal models into production.
  • Optimize inference latency, throughput, memory usage, GPU utilization, and cost efficiency.
  • Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes.
  • Collaborate with platform, infrastructure, product, and customer-facing teams and contribute to open-source inference projects when appropriate.

Requirements

  • 5+ years of experience writing reliable, well-tested production code.
  • Strong Python skills and experience designing production systems.
  • Hands-on experience with PyTorch and deploying machine learning models.
  • Experience with Linux, Docker, and Kubernetes.
  • Experience in distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling.
  • Ability to debug complex problems across software, infrastructure, and hardware, with strong communication and collaboration skills.

Nice to have

  • Experience with vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or similar inference frameworks.
  • Experience running GPU workloads in production and optimizing model performance.
  • Knowledge of quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, or LoRA serving.
  • Experience with CUDA, Triton, TensorRT, distributed inference, multi-GPU systems, scheduling, or autoscaling.
  • Contributions to open-source ML, inference, or infrastructure projects.

Culture & Benefits

  • Employment is available only under a labor agreement.
  • Flexible working hours and hybrid or remote options may be available depending on the role.
  • Private medical insurance, extra paid vacation, and sick leave days may be provided depending on location.
  • Language courses, support for important life events, team sports, and social activities.
  • Modern offices with snacks, drinks, and entertainment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →