Назад
Company hidden
2 дня назад

Site Reliability Engineer

Формат работы
remote (только Argentina/Brazil/Canada)
Тип работы
fulltime
Английский
b2
Страна
UK/Singapore/US +6 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Kubernetes/AWS): Maintaining and improving the global Tyk Cloud platform with an accent on reliability engineering, infrastructure automation, observability, and incident management. Focus on operating multi-region and multi-cloud Kubernetes environments, building metrics and dashboards, automating operations, and resolving complex production reliability issues.

Location: Remote; listed hiring locations include Argentina, Brazil, Canada, Colombia, Costa Rica, and Mexico. Remote working from anywhere in the world is offered. On-call coverage is required from 16:00–04:00 UTC.

Company

hirify.global provides an API Management platform that helps organizations connect systems and services through its cloud and B2B products.

What you will do

  • Maintain and improve the global hirify.global Cloud platform within defined service-level objectives.
  • Identify and resolve reliability issues with the engineering squad.
  • Expand the platform’s multi-region and multi-cloud capabilities.
  • Build metrics, dashboards, monitoring, and logging systems.
  • Participate in on-call rotation, incident management, post-incident analysis, and penetration-testing support.
  • Automate operational tasks, document SRE processes, and improve operational efficiency and running costs.

Requirements

  • Experience launching and operating production-scale Kubernetes clusters and containerized infrastructure.
  • Advanced AWS/EKS and Linux administration experience, plus infrastructure design on AWS or other cloud providers.
  • Proficiency with Terraform, infrastructure as code, and Helm.
  • Experience operating MongoDB or other document databases, Redis or other key-value stores, and distributed software.
  • Experience with Prometheus, Grafana, Thanos, logging systems, networking concepts, and common networking protocols.
  • Strong collaboration skills and willingness to participate in on-call rotation from 16:00–04:00 UTC.

Nice to have

  • Experience with GCP or Azure, bare-metal infrastructure, API management, or large-scale distributed storage.
  • Familiarity with Rancher and CKA, CKAD, or CKS certifications.
  • Experience creating and delivering production software in Go.

Culture & Benefits

  • Unlimited paid holiday and flexible working hours.
  • Remote-first work with a distributed international team.
  • Employee share scheme.
  • Generous maternity and paternity leave.
  • Company retreats and a culture centered on autonomy, responsibility, inclusion, experimentation, and continuous improvement.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →