Назад
Company hidden
5 дней назад

Senior Site Reliability Engineer (AWS/Kubernetes)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Brazil
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AWS/Kubernetes): Building and operating reliable, scalable infrastructure for a multi-tenant SaaS platform with an accent on automation, observability, security, and production readiness. Focus on designing self-healing systems, improving resilience through incident analysis, and enabling follow-the-sun operations across distributed engineering teams.

Location: Curitiba, Brazil

Company

hirify.global provides a cloud platform for property and casualty insurers, combining core insurance, digital, analytics, and AI capabilities.

What you will do

  • Design, build, and operate highly reliable, scalable infrastructure for a multi-tenant SaaS platform.
  • Automate deployment, provisioning, and operational workflows across cloud infrastructure and applications.
  • Build internal tools, services, frameworks, and observability systems covering metrics, logging, tracing, and dashboards.
  • Define SLOs, investigate incidents, lead root cause analysis and blameless postmortems, and reduce operational toil.
  • Partner with development teams on availability, performance, scalability, security, and production readiness.
  • Mentor engineers and create documentation, runbooks, and training materials.

Requirements

  • Strong programming skills in Python or Go.
  • Deep experience with AWS and production systems operating at scale.
  • Hands-on expertise with Kubernetes, including EKS, Docker, Helm, CNI, Ingress, and Kubernetes primitives.
  • Experience with Infrastructure as Code using Terraform, Terragrunt, or similar tools.
  • Solid Linux and networking fundamentals, plus experience with observability platforms such as Datadog, Prometheus, OpenTelemetry, or CloudWatch.
  • Experience with CI/CD, GitOps, incident management, microservices production support, SSO, SAML, OAuth, and identity providers.

Nice to have

  • Java/Spring Boot experience.
  • Experience with Kafka, SQS, Aurora, or RDS.
  • Exposure to KubeVela, Crossplane, AWS or Kubernetes certifications, or open-source contributions.
  • Bachelor’s degree in Computer Science or a related field, or equivalent experience.

Culture & Benefits

  • Work on a mission-critical global platform used by leading insurers.
  • Solve complex real-world infrastructure and reliability problems at scale.
  • Collaborate with distributed engineering teams in a high-impact environment.
  • Participate in a 24x7 follow-the-sun on-call rotation for critical production systems.
  • Use AI and data-driven insights to improve engineering productivity and outcomes.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →