обновлено 11 дней назад
Technical Operations Lead (AI Cloud Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Technical Operations Lead (AI Cloud Platform): Operating and supporting a large-scale private AI cloud platform with an accent on reliability, incident management, disaster recovery, and platform performance. Focus on coordinating major incidents, managing Kubernetes and OpenStack environments, testing backup and recovery capabilities, and driving continuous operational improvement.
Location: Athens, Attica, Greece; Workplace: Hybrid
Company
Quento Technologies S.A. is the ICT arm of , delivering solutions across AI, digital engineering, cloud, and cybersecurity.
What you will do
- Lead day-to-day operations and technical support for a large-scale private AI cloud platform.
- Ensure platform availability, performance, reliability, capacity, and operational readiness.
- Own incident, problem, change management, major incident response, and technical escalation processes.
- Maintain operational procedures, runbooks, and support models while coordinating maintenance, upgrades, and vendor support.
- Lead backup, restore, and disaster recovery testing and drive root cause analysis and continuous improvement.
- Ensure operational activities comply with regulatory requirements and the Group Anti-Bribery and Corruption Policy.
Requirements
- 5+ years of experience in infrastructure, cloud, or platform operations and 3+ years in a technical lead or operations leadership role.
- Experience operating enterprise-scale or mission-critical platforms and supporting 24x7 operations under SLAs.
- Strong experience with Linux, Kubernetes, private cloud environments, backup, disaster recovery, and business continuity.
- Technical knowledge of Canonical OpenStack, Canonical Kubernetes, MAAS, Juju, Ceph, Ubuntu Server, Prometheus, Grafana, Elasticsearch, OpenSearch, Loki, Terraform, Ansible, Azure DevOps, Keycloak, Vault, PAM, SIEM solutions, and NVIDIA GPU infrastructure.
- Strong stakeholder management and vendor coordination skills.
Nice to have
- Experience with AI or GPU platform operations.
- Knowledge of Site Reliability Engineering practices.
- ITIL certification or equivalent practical experience.
- Familiarity with ISO 27001 or ISO 22301 and regulated environments.
Culture & Benefits
- Competitive compensation, restaurant card, and annual bonus programs.
- Modern IT equipment, mobile and data plan, and home equipment benefits.
- Private health insurance, occupational doctor, workplace counselor, gym, and wellness facilities.
- Hybrid working model, career development tools, mentoring, coaching, and a personalized annual learning plan.
- Modern facilities with free beverages, indoor parking, company bus, and employee wellbeing, ESG, and volunteering activities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Engineering Manager (AI)
12 дней назад
Cloud Engineer (Microsoft Azure)
Toughbyte
3 часа назад
SRE / DevOps Engineer (GCP, Kubernetes)
5 500€
11 дней назад
Team Lead Development (Node.js/TypeScript)
14 дней назад
Chapter Lead Feature Engineer (Enterprise Content & Automation)
Kiwitaxi
3 часа назад
Backend TechLead
4 500$