Назад
Company hidden
12 часов назад

GPU Platform Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
Kazakhstan
Релокация
Kazakhstan
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

GPU Platform Engineer (AI): Owning and optimizing the GPU infrastructure for generative video and image products with an accent on high-utilization inference fleets and large-scale training clusters. Focus on building robust, GitOps-driven Kubernetes environments, maximizing GPU throughput, and automating capacity management at the frontier of AI-native experiences.

Location: On-site in Almaty, Kazakhstan (relocation provided)

Company

hirify.global is a rapidly scaling generative AI company powering creative tools for millions of users and Fortune 500 brands.

What you will do

  • Optimize distributed training clusters for maximum throughput using NCCL tuning and topology-aware scheduling.
  • Manage the lifecycle of a multi-provider GPU fleet using Talos Linux and Sidero Omni.
  • Implement and maintain KEDA-driven autoscaling for inference workloads to ensure cost-efficiency.
  • Enforce GitOps practices across all clusters using ArgoCD, Helm, and Terraform.
  • Develop observability dashboards and SLOs to monitor GPU utilization, queue latency, and cost per generation.
  • Collaborate with ML engineers to streamline model-serving rollouts and runtime performance.

Requirements

  • 3+ years of experience running production Kubernetes as an SRE, Platform, or MLOps engineer.
  • Hands-on experience with distributed training operations, including NCCL and high-speed interconnects.
  • Proficiency with bare-metal Kubernetes and immutable OS setups like Talos.
  • Deep knowledge of the NVIDIA stack, including drivers, GPU Operator, and DCGM metrics.
  • Strong GitOps fluency and experience with infrastructure-as-code tools like Terraform.
  • Proficiency in Python or Go for automation and scripting.

Nice to have

  • C++ and CUDA programming skills for custom kernel development and profiling.
  • Experience with Sidero Omni and multi-provider GPU clouds like Nebius or CoreWeave.
  • Knowledge of inference runtimes such as Triton, vLLM, or TensorRT.
  • Experience with FinOps for GPU fleets and internal AIOps tooling.

Culture & Benefits

  • Competitive base salary paid in USD.
  • Equity participation in the company’s stock option program.
  • Relocation assistance to Almaty, Kazakhstan.
  • Opportunity to work at the absolute frontier of generative AI technology.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →