2 дня назад
Director of Platform Engineering (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Director of Platform Engineering (AI/Kubernetes): Leading the reliability and continuous improvement of an Analytics SaaS observability platform that collects, stores, and analyses data from banks and financial institutions with an accent on Kubernetes operations, cloud-native infrastructure, and service resilience. Focus on automating operational workflows, applying Agentic AI to incident management, and maintaining secure, auditable practices for regulated customers.
Location: London, UK; hybrid working with typically 2 days per week in the office
Company
develops automated IT observability solutions for critical and regulated industries, including financial institutions, through an Analytics SaaS platform.
What you will do
- Lead the day-to-day management, reliability, and continuous improvement of the Analytics SaaS platform.
- Provide hands-on technical leadership for platform and SaaS engineering teams, including coaching, standards, and accountability for service outcomes.
- Work directly on Kubernetes troubleshooting, production incidents, deployment automation, observability, capacity management, and service resilience.
- Reduce operational toil through workflow automation, self-service tooling, standardized runbooks, and self-healing capabilities.
- Explore Agentic AI and intelligent automation for incident triage, runbook execution, knowledge retrieval, change preparation, and service management.
- Lead the Forward Deployed Engineering function and coordinate with Engineering, Product, Security, and customer-facing teams.
Requirements
- Experience leading SaaS Hosting or operational teams in a hands-on technical leadership role, ideally in a high-availability, enterprise, or regulated environment.
- Deep practical Kubernetes operations knowledge, including deployments, upgrades, networking, storage, ingress, secrets, autoscaling, and cluster reliability.
- Strong cloud operations experience with AWS, Azure, or GCP, plus Infrastructure as Code, CI/CD, GitOps, or similar automation approaches.
- Experience improving reliability and reducing operational toil through automation, workflow redesign, self-service capabilities, runbooks, and observability.
- Understanding of secure and auditable Agentic AI or intelligent automation, SaaS security, access control, change management, vulnerability management, and operational resilience.
- Experience leading or closely supporting Forward Deployed Engineering, Solutions Engineering, Customer Engineering, or similar customer-embedded teams.
Nice to have
- Experience with managed Kubernetes services such as Amazon EKS, Azure AKS, or Google GKE.
- Knowledge of GitOps tools such as Argo CD or Flux.
- Experience with ISO 27001, SOC 2, SRE practices, service-level objectives, capacity planning, or production readiness reviews.
- Experience applying AI, AIOps, or intelligent automation to incident triage, anomaly detection, runbook automation, or root-cause support.
- Background in financial services, enterprise technology, DevOps, SRE, Cloud Operations, or Platform Engineering.
Culture & Benefits
- Flexible hybrid working in an inclusive and supportive environment.
- Health and dental insurance for employees and dependants.
- Pension, life assurance, income protection, travel insurance, and enhanced parental leave.
- Employee Assistance Programme, training reimbursement, referral bonus, and buy-and-sell holiday options.
- Work supporting critical technology used by global customers in demanding and regulated industries.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Platform Engineering Manager (AI)
4 дня назад
Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)
9 дней назад
Director, Site Reliability Engineering
Nscale
6 дней назад
Deployment Engineering Director, Systems Engineering
5 дней назад
Senior Production Engineer (Fintech)
3 дня назад