9 дней назад
T-Hub - AIOps Engineer - AI Infrastructure & Orchestration
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
T-Hub - AIOps Engineer - AI Infrastructure & Orchestration (vLLM/Kubernetes): Designing and operating scalable AI inference services on OpenShift/Kubernetes with an accent on GPU orchestration, model lifecycle automation, observability, and secure API delivery. Focus on optimizing LLM serving performance, building token-usage monitoring and chargeback capabilities, and troubleshooting distributed AI infrastructure.
Location: Not specified
Contract: Employment contract. Working mode: Full time.
Company
is a telecommunications company developing connectivity and technology solutions spanning 5G, IoT, and AI.
What you will do
- Design, deploy, and maintain vLLM inference services on OpenShift and Kubernetes running on bare-metal GPU infrastructure.
- Manage NVIDIA GPU partitioning and allocation across multiple models and tenants, optimizing utilization and serving performance.
- Automate model onboarding, versioning, deployment, hot-swapping, and rollback from private registries and S3.
- Build observability solutions with metrics, logging, tracing, Grafana dashboards, Prometheus, and Alertmanager.
- Implement usage tracking for token consumption, quota management, and chargeback or showback requirements.
- Provide API gateway, security, audit logging, incident response, root cause analysis, and continuous platform improvement in collaboration with AI, platform, security, and infrastructure teams.
Requirements
- 5+ years of experience in DevOps, SRE, platform engineering, or infrastructure operations.
- At least 2 years of hands-on experience with MLOps, AI infrastructure, or LLM platforms.
- Production experience administering Kubernetes and OpenShift and operating vLLM-based inference platforms.
- Knowledge of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization.
- Experience with NVIDIA GPUs, CUDA drivers, NVIDIA Container Toolkit, GPU troubleshooting, Prometheus, Grafana, OpenTelemetry, and ELK Stack.
- Strong Python, Bash, and Linux administration skills, plus familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.
Culture & Benefits
- Work with modern technologies including 5G, IoT, and AI.
- Collaborate across AI engineering, platform engineering, security, and infrastructure functions.
- Contribute to innovation in the telecommunications and connectivity domain.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →