6 дней назад
Python Inference Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Python Inference Engineer (AI): Building and improving the inference layer of Gcore's AI platform with an accent on model deployment, inference frameworks, GPU utilization, and performance optimization. Focus on integrating language and multimodal models, debugging across software and infrastructure layers, and improving latency, throughput, memory use, reliability, and cost efficiency.
Location: Poland, Serbia, or Cyprus; hybrid or remote options may be available depending on the role. Work from anywhere in the world is available for up to 45 days per year.
Company
provides infrastructure and software for AI, cloud, network, and security workloads through a global edge and cloud platform.
What you will do
- Build and improve the inference layer of the Inference platform.
- Integrate and operate inference frameworks including vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM.
- Bring new language and multimodal models into production.
- Optimize inference latency, throughput, memory usage, GPU utilization, and cost efficiency.
- Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes.
- Collaborate with platform, infrastructure, product, and customer-facing teams, and contribute to open-source inference projects when appropriate.
Requirements
- 5+ years of experience writing reliable, well-tested production code.
- Strong Python skills and experience designing production systems.
- Hands-on experience with PyTorch and deploying machine learning models.
- Experience with Linux, Docker, Kubernetes, and at least one of distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling.
- Ability to debug complex issues across software, infrastructure, and hardware, with strong communication and collaboration skills.
- This position is available only under an employment (labor) agreement.
Nice to have
- Experience with vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or similar inference frameworks.
- Experience running GPU workloads in production, including distributed inference, multi-GPU systems, scheduling, or autoscaling.
- Knowledge of quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, or LoRA serving.
- Experience with CUDA, Triton, TensorRT, profiling, or open-source ML and infrastructure projects.
Culture & Benefits
- Flexible working hours with hybrid or remote options depending on the role.
- Private medical insurance for employees and families, where applicable.
- Extra paid vacation and sick leave days, depending on location.
- Support for important life events, language courses, team sports, and social activities.
- Modern offices with snacks, drinks, and entertainment, where applicable.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →