2 дня назад
DevOps Engineer (AI Inference)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
DevOps Engineer (AI Inference) (Kubernetes/GPU infrastructure): Designing and operating on-premises infrastructure and services for scalable, secure AI inference workloads with an accent on GPU scheduling, model deployment pipelines, observability, and production Kubernetes. Focus on integrating inference runtimes, troubleshooting Linux and networking issues, automating operations with Python, Go, or Bash, and testing performance at scale.
Location: Poland, Serbia, Georgia, Cyprus, or Germany; hybrid or remote options are available depending on the role. Work from anywhere in the world is available for up to 45 days per year.
Company
provides infrastructure and software solutions for AI, cloud, network, and security, operating edge locations, cloud regions, and GPU infrastructure worldwide.
What you will do
- Design, develop, and maintain on-premises infrastructure for AI inference workloads, including GPU scheduling, model deployment pipelines, and data access patterns.
- Build and manage monitoring and observability systems, including dashboards, alerts, and runbooks for model health and system performance.
- Collaborate with ML engineers and platform teams on AI workload architecture and inference runtime integration.
- Test AI inference platform performance at scale and maintain reliable production services.
Requirements
- Strong understanding of Kubernetes architecture, including CNI, CSI, operators, ingress/gateway, and control plane components.
- Hands-on experience operating and troubleshooting production Kubernetes clusters.
- Strong Linux and networking troubleshooting skills covering DNS, routing, firewalling, TLS, MTU, connectivity, and performance.
- Ability to develop automation and operational tooling with Python, Go, or Bash.
- Experience with Terraform, Ansible, monitoring and alerting tools such as VictoriaMetrics or Grafana, Git workflows, and CI/CD pipelines.
- Employment is available only under a labor agreement.
Nice to have
- Familiarity with Cluster API, Slurm, Argo CD, GitOps, Helm, or Helmfile.
- Experience with managed platforms, PaaS, cloud services, bare metal, GPU, or HPC environments.
- Knowledge of the NVIDIA GPU stack, RDMA/InfiniBand, OpenStack, or Kubernetes operators and controllers.
Culture & Benefits
- Flexible working hours and hybrid or remote options depending on the role.
- Private medical insurance for employees and families, where available.
- Extra paid vacation and sick leave days, depending on location.
- Language courses, support for important life events, team sports, and social activities.
- Modern offices with snacks, drinks, and entertainment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →