6 дней назад
Senior Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI): Building and operating resilient, scalable platforms and services that support critical applications and AI-powered features with an accent on AWS operations, Terraform automation, observability, and incident response. Focus on designing reliability workflows, testing recovery capabilities, automating failovers and rollbacks, and enabling engineering teams to operate secure self-service platforms.
Location: London Wall, London, United Kingdom
Company
provides information, analytics, publishing, research, healthcare, and learning solutions that support science and clinical practice.
What you will do
- Create monitoring queries, service-level baselines, observability dashboards, SLOs, and error budgets.
- Improve reliability, scalability, performance, and recovery of critical platforms and services.
- Implement production automation, including incident-response scripts, failovers, rollbacks, and CI/CD workflows.
- Participate in incident response, on-call support, post-mortems, root-cause analyses, and disaster-recovery testing.
- Support the deployment, monitoring, reliability, and security of services integrating AI tools.
- Enable engineering teams through platform migration, documentation, consultancy, and secure self-service practices.
Requirements
- Advanced Terraform expertise, including modules, providers, state management, drift detection, safe refactoring, and remote state.
- Hands-on experience operating production multi-account and multi-region AWS environments, including ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
- Strong experience with GitHub Actions CI/CD, reusable workflows, OIDC authentication, approval gates, runners, and deployment pipelines.
- Strong knowledge of Docker, ECS Fargate, containers, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
- Advanced DevOps experience across monitoring, networking, cloud storage, orchestration, configuration management, Linux, Git, Bash, and Python automation.
- Experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices.
Culture & Benefits
- Flexible working hours and emphasis on work-life balance.
- Comprehensive pension plan and generous vacation entitlement.
- Sabbatical leave and family-related leave, including maternity, paternity, adoption, and family care leave.
- Study assistance, wellbeing initiatives, internal communities, employee discounts, and an employee assistance program.
- Recruitment introduction reward and country-specific benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Principal Site Reliability Engineer (AWS/Terraform)
107 000GBP
7 дней назад
Senior DevOps Engineer
75 000 - 85 000$
10 дней назад
Lead Site Reliability Engineer (Kubernetes)
11 дней назад
Site Reliability Engineer (Azure)
6 дней назад
Senior Site Reliability Engineer (AI)
75 000 - 85 000$
11 дней назад