4 дня назад
NOC Engineer / SRE
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
NOC Engineer / SRE (Linux, Kubernetes, Cloud, Observability): Ensuring 24/7 service reliability through incident response, monitoring, operational automation, and production infrastructure support with an accent on observability, alerting, and reducing operational toil. Focus on engineering self-healing workflows, improving MTTD and MTTR, and supporting reliable Linux, cloud, and Kubernetes environments.
Location: United Kingdom - Remote
Company
develops AI, cloud, and digital software used by global businesses for customer experience, financial crime prevention, and public safety.
What you will do
- Participate as a primary or escalation responder in a 24/7 on-call rotation.
- Lead and support major incident response, including triage, mitigation, resolution, and blameless post-incident reviews.
- Coordinate incident response across Engineering, Infrastructure, Security, and Product teams.
- Own service health monitoring and design alerting strategies aligned with SLIs and SLOs.
- Build dashboards, reduce alert fatigue, and improve observability across infrastructure, applications, and dependencies.
- Automate operational tasks, implement self-healing and auto-remediation, and support production releases.
Requirements
- Strong Linux systems administration and experience with incident management and production support.
- Experience working in 24/7 NOC or production operations environments.
- Experience with cloud infrastructure, preferably AWS, plus Docker and Kubernetes.
- Scripting or programming experience in Python, Bash, Go, or similar technologies.
- Knowledge of monitoring and alerting platforms such as Grafana, Prometheus, Datadog, Splunk, or CloudWatch.
- Understanding of networking fundamentals, including DNS, TCP/IP, and load balancing.
hirify.global-to-have"> to have
- Experience defining or operating to SLOs and SLIs.
- Experience migrating from a traditional NOC to an SRE model.
- Infrastructure as Code experience with Terraform, Ansible, or similar tools.
- Exposure to security, compliance, or regulated environments.
Culture & Benefits
- Remote work in the United Kingdom.
- Individual contributor role reporting to the Manager, Network Operations.
- Work in a 24/7 operational environment focused on reliability, automation, and continuous improvement.
- Opportunity to support software used by more than 25,000 global businesses.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
7 дней назад
Staff Service Reliability and Operational Intelligence Engineer (AI Ops)
Replit
7 дней назад
Site Reliability Engineer
210 000 - 275 000$
6 дней назад
Site Reliability Engineering Manager (AWS/Kubernetes)
205 000 - 255 000$
10 дней назад
Senior Site Reliability Engineer (AWS/Kubernetes)
60 000 - 85 500€
Replit
7 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$