Назад
Company hidden
5 дней назад

Technical Operations Lead (AI Cloud Platform)

Формат работы
hybrid
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
Belgium/Greece/Luxembourg
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Operations Lead (AI Cloud Platform) (AI, Kubernetes, OpenStack): Operating and improving a large-scale private AI cloud platform with an accent on reliability, incident management, disaster recovery, and enterprise infrastructure. Focus on coordinating major incident response, managing capacity and performance, and driving root cause analysis across Kubernetes, OpenStack, Ceph, GPU infrastructure, and observability systems.

Location: Athens, Attica, Greece; Workplace: Hybrid

Company

Quento Technologies S.A. is the ICT arm of the hirify.global, delivering AI, digital engineering, cloud, and cybersecurity solutions.

What you will do

  • Lead day-to-day operations and technical support for a large-scale private AI cloud platform.
  • Ensure platform availability, performance, reliability, capacity, and operational readiness.
  • Own incident, problem, change, escalation, and major incident management processes.
  • Maintain operational procedures, runbooks, and support models.
  • Lead backup, restore, disaster recovery testing, maintenance, upgrades, and vendor coordination.
  • Drive root cause analysis, continuous improvement, and compliance with regulatory and anti-bribery requirements.

Requirements

  • 5+ years of experience in infrastructure, cloud, or platform operations and 3+ years in a technical lead or operations leadership role.
  • Experience operating enterprise-scale or mission-critical platforms and supporting 24x7 operations with major incident and SLA management.
  • Strong experience with Linux, Kubernetes, private cloud environments, backup, disaster recovery, and business continuity.
  • Technical knowledge of Canonical OpenStack, Canonical Kubernetes, MAAS, Juju, Ceph, Ubuntu Server, Prometheus, Grafana, Elasticsearch, OpenSearch, Loki, Terraform, and Ansible.
  • Experience with Azure DevOps or equivalent, Keycloak, Vault, PAM, SIEM solutions, and NVIDIA GPU infrastructure.
  • Strong stakeholder management and vendor coordination skills.

Nice to have

  • Experience with AI or GPU platform operations.
  • Knowledge of Site Reliability Engineering practices.
  • ITIL certification or equivalent practical experience.
  • Familiarity with ISO 27001 or ISO 22301 and regulated environments.

Culture & Benefits

  • Hybrid working model with home equipment benefits.
  • Competitive compensation, restaurant card, and annual bonus programs.
  • Private health insurance, occupational doctor, and workplace counselor.
  • IT equipment, mobile and data plan, modern facilities, parking, and company bus.
  • Career development, mentoring, coaching, and a personalized annual learning plan.
  • Wellbeing, ESG, volunteering, gym, and wellness activities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →