5 дней назад
MLOps Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
MLOps Engineer (AI) (Kubernetes/GPU/LLM): Building and operating Kubernetes infrastructure for production GPU workloads and scalable LLM inference services with an accent on deployment automation, observability, and resource optimization. Focus on designing reliable serving platforms, improving GPU utilization and traffic routing, and handling production incidents across stateful ML infrastructure.
Location: Yerevan, Armenia
Company
operates production machine learning and LLM workloads supported by GPU-based infrastructure.
What you will do
- Build and operate Kubernetes infrastructure for production GPU and machine learning workloads.
- Design scalable deployment patterns and reliable serving infrastructure for LLM inference and AI services.
- Own CI/CD pipelines, Helm charts, deployment automation, and infrastructure configuration.
- Design observability, metrics, SLOs, dashboards, and alerting with Prometheus, Grafana, and Alertmanager.
- Optimize GPU utilization, workload scheduling, autoscaling, batching, traffic routing, and container performance.
- Operate stateful services including Kafka, Redis, ClickHouse, and PostgreSQL; participate in on-call, incident response, root cause analysis, and corrective actions.
Requirements
- 3+ years of experience in MLOps, ML platform engineering, DevOps, or SRE, including 2+ years running GPU workloads in production.
- Strong Kubernetes experience with GPU device plugins, node selectors and taints, resource management, and HPA/KEDA autoscaling.
- Experience authoring and customizing Helm charts and owning CI/CD with GitHub Actions on self-hosted runners.
- Production observability experience with Prometheus, Grafana, and Alertmanager, including metrics, SLOs, and alert rules.
- Experience operating Kafka, Redis, ClickHouse, and PostgreSQL on Kubernetes, including backups and restore drills.
- Strong Docker, shell, and Python skills, plus production on-call and incident response experience.
Nice to have
- Bare-metal or on-premises GPU cluster setup, GPU Operator, driver lifecycle, NVLink, InfiniBand, local registries, and object storage.
- Cloud GPU infrastructure experience with Nebius, AWS, GCP, or Azure.
- Experience with Envoy Gateway, Gateway API, NVIDIA Dynamo, AIPerf, GenAI-Perf, or LLM-aware request routing.
- Knowledge of secrets management, RBAC, NetworkPolicy, image signing, scanning, and supply-chain security.
- Experience with MLflow, Kubeflow, Metaflow, Argo CD, Flux, MIG, or GPU time-slicing.
Culture & Benefits
- Collaboration with ML Engineering, Platform, and Infrastructure teams.
- Ownership of production infrastructure, reliability, scalability, and observability.
- Opportunity to evaluate technologies across GPU infrastructure, Kubernetes, LLM serving, and ML platform engineering.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →