Назад
2 дня назад

Senior Site Reliability Engineer (Kubernetes)

Формат работы
remote (Global)
Тип работы
fulltime
Грейд
senior
Страна
Brazil
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Senior Site Reliability Engineer (Kubernetes/DevOps): Building reliable, observable, and self-healing infrastructure at scale with an accent on automation, cloud-native reliability, monitoring, and incident response. Focus on designing resilient systems, improving observability and alerting, leading post-incident reviews, and adopting SLO and SLI practices.

Senior Site Reliability Engineer

Company

Latitude

Conditions

3 days agoSenior Anywhere Remote Contract Devops Jobs by Latitude

Latitude Latitude.sh provides global bare metal cloud infrastructure with automated deployment and management through a unified API and dashboard. Company intelligence Brazil Projects Latitude Filesystem storage Decentralized File Storage Accelerate Compute Network Launchpad Developer Tooling Metal Compute Network D Databases Off-Chain Data APIs and Services C Cloud Gateway Developer Tooling About Latitude Latitude is a global Bare Metal Cloud provider with an automated platform that enables instant deployment and easy management of dedicated, high-performance, low-latency servers around the world, optimized for the digital economy. View jobs by Latitude

Skills

Ansible Bash Ci/Cd Container Orchestration Elk Git Go Grafana Incident Management Kubernetes Linux Loki Observability Prometheus Python Root-Cause Analysis Ruby Sli Slo Terraform Unix

Candidate Availability

Remote · Required Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build reliable, observable, and self-healing infrastructure at scale. You will automate operational tasks and incident response, improve monitoring and alerting, collaborate on resilient system designs, participate in on-call rotations, lead post-incident reviews, and document operational processes and runbooks.

Requirements

  • Advanced knowledge of Linux/Unix systems in production environments
  • Experience with Kubernetes and container orchestration
  • Proficiency with Terraform and Ansible
  • Experience with Prometheus, Grafana, Loki, or ELK
  • Familiarity with Bash, Python, Go, or Ruby
  • Working knowledge of Git and CI/CD pipelines
  • Understanding of incident management and root cause analysis
  • Knowledge of cloud-native reliability and security best practices

Responsibilities

  • Improve platform reliability and performance
  • Design, build, and maintain tools that automate operational tasks and incident response
  • Implement and improve monitoring, alerting, and tracing solutions
  • Collaborate on scalable and resilient system designs
  • Participate in on-call rotations
  • Lead post-incident reviews
  • Develop and document operational processes and runbooks
  • Contribute to SLO, SLI, and reliability-metric adoption

Benefits

  • Paid Time Off
  • Wellhub
  • Annual bonus based on company and team performance
  • Flexible work hours

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -