Назад
Company hidden
обновлено 7 дней назад

Senior Site Reliability Engineer (Kubernetes/AWS)

Формат работы
remote (Global)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/Spain/Ireland +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Kubernetes/AWS): Operating and improving Develocity instances, supporting services, and cloud infrastructure with an accent on reliability, observability, automation, and incident response. Focus on building self-healing deployments, managing disaster recovery, troubleshooting across the application and infrastructure stack, and optimizing performance and costs.

Location: Remote, must be located in the GMT timezone: UK, Ireland, Portugal, or the Canary Islands

Company

hirify.global builds Develocity, a SaaS toolchain observability and intelligence platform with build acceleration, deep observability, and AI-powered capabilities.

What you will do

  • Operate and maintain Develocity instances, artifact registries, and supporting services.
  • Participate in a 24/7 on-call rotation, lead incident response, troubleshoot issues across the application and infrastructure stack, and communicate with customers during incidents and maintenance windows.
  • Automate deployments, upgrades, monitoring, self-healing, recovery, backups, and disaster-recovery processes.
  • Build and maintain observability through logging, metrics, tracing, and alerting.
  • Work with engineering and Cloud Platform teams to build reliability into features and improve internal platform tooling.
  • Optimize performance, resource usage, costs, and SaaS operations as the platform grows.

Requirements

  • 5+ years of experience in SRE, DevOps, or an equivalent role operating production services at scale.
  • Strong production Kubernetes experience and cloud infrastructure expertise, preferably with AWS, including EKS, RDS, S3, and EC2.
  • Proficiency with Prometheus, Grafana, Terraform, Python, and Bash.
  • Experience with incident management, 24/7 on-call rotations, SRE practices, SLAs, and SLOs.
  • Strong written and verbal English communication skills required.
  • Must be located in the GMT timezone.

Nice to have

  • Experience operating SaaS platforms at scale or establishing SRE practices in new or growing teams.
  • Familiarity with Develocity and JVM languages such as Java or Kotlin.
  • Experience with disaster recovery planning, execution, and customer-facing incident communication.

Culture & Benefits

  • Remote-first, work-from-home environment with asynchronous communication and written documentation.
  • Founding role in a new SRE team with ownership of operational practices and production systems.
  • Culture focused on automation over heroics, continuous learning, and clear ownership of outcomes.
  • In-person annual offsites and team meetings.
  • Competitive salary and equity grants.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →