4 дня назад
Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff) (Kubernetes/Infrastructure as Code): Building and operating reliable, scalable production infrastructure for GitLab's user-facing services with an accent on automation, Kubernetes operations, observability, and safe delivery. Focus on designing infrastructure tooling, responding to incidents, improving SLO-driven reliability, and solving systemic infrastructure challenges across teams.
Location: Remote for candidates based in the United Kingdom only
Company
is a DevSecOps platform that helps organizations improve developer productivity, operational efficiency, security, compliance, and digital transformation.
What you will do
- Keep 's user-facing services and production systems reliable, scalable, and efficient.
- Build automation, infrastructure tooling, and infrastructure-as-code workflows that reduce toil and manual work.
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
- Write and maintain infrastructure as code and deliver changes safely through CI/CD and GitOps.
- Participate in on-call rotations, alert triage, incident response, runbook improvement, and post-incident reviews.
- Improve observability through metrics, logs, alerting, SLOs, documentation, and repeatable operational practices.
Requirements
- Experience maintaining reliable production systems through software engineering and operational practices.
- Experience building new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators, controllers, or production services.
- Ability to read, debug, and reason about code; most teams use Go, with some using Ruby.
- Experience with infrastructure as code, Kubernetes and its ecosystem, and at least one major cloud provider: AWS or GCP.
- Familiarity with observability, including metrics, logging, alerting, SLOs or SLIs, and data-informed operational decisions.
- Comfort with on-call and incident response, strong written communication, async collaboration, and manager-of-one ownership.
Nice to have
- Experience using AI and automation to reduce toil and improve individual and team workflows.
- Experience influencing reliability across multiple teams or setting technical direction at organizational scale.
Culture & Benefits
- Remote-first, globally distributed, asynchronous organization.
- Flexible paid time off and parental leave.
- Health, financial, and well-being benefits.
- Equity compensation and an employee stock purchase plan.
- Growth and development fund and team member resource groups.
Hiring process
- Recruiter screen followed by a shared core technical assessment.
- Hiring manager, peer technical, and skip-level interviews.
- Level and team placement are calibrated throughout the process based on experience, interview results, and hiring needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Staff Site Reliability Engineer (GCP/Kubernetes)
6 дней назад
Site Reliability Engineer (SRE)
8 дней назад
Senior Site Reliability Engineer I (Kubernetes)
8 дней назад
Senior Site Reliability Engineer (Kubernetes)
11 дней назад
Principal SRE, Infrastructure & Platform
DeepL
6 дней назад