Назад
Company hidden
1 день назад

Site Reliability Engineer (Kubernetes)

123 000 - 150 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Kubernetes): Building, operating, and scaling a multi-region SaaS platform across AWS, GCP, and Azure with an accent on Kubernetes infrastructure, service mesh architectures, and high-availability data layers. Focus on automating GitOps deployments, improving observability and incident response, and solving complex reliability, scalability, and resilience challenges in a 24/7 production environment.

Location: Remote, Washington, United States

Salary: $123,000–$150,000 per year; compensation varies by location, skills, experience, and role level. Benefits may vary by location.

Company

Develops API and AI connectivity technologies and a unified platform for securing, managing, governing, and monetizing API and AI model traffic.

What you will do

  • Operate and scale a multi-region, multi-cloud SaaS platform across AWS, GCP, and Azure.
  • Build Kubernetes infrastructure and deployment workflows with Terraform, Terragrunt, Helm, and ArgoCD.
  • Design and optimize highly available, low-latency data and caching layers using PostgreSQL, Redis, ClickHouse, and Druid.
  • Operate API gateway and service mesh environments supporting hybrid and distributed architectures.
  • Develop CI/CD pipelines and GitOps workflows, and improve observability with Datadog, Prometheus, Grafana, and Thanos.
  • Participate in a 24/7 on-call rotation, lead scaling initiatives, and improve incident response, postmortems, and operational playbooks.

Requirements

  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • Experience managing enterprise-scale SaaS or PaaS systems in secure, multi-region, and multi-tenant environments.
  • Deep Kubernetes expertise, including cluster and networking troubleshooting, fault tolerance, and scalability design.
  • Strong proficiency with Infrastructure as Code, especially Terraform or Terragrunt, plus CI/CD and GitOps workflows.
  • Programming experience with Go, Python, or Bash; solid knowledge of Linux/Unix, DNS, TLS/SSL, HTTP, load balancers, and distributed systems.
  • Experience with API gateways, service meshes, Kafka, observability platforms, and 24/7/365 production support.

Nice to have

  • Experience with hirify.global Gateway, hirify.global Mesh, ClickHouse, Druid, PostgreSQL, or Redis in multi-region environments.
  • Knowledge of AWS networking, Azure VNet, or GCP NCC.
  • Experience with disaster recovery, resiliency testing, and compliance-driven reliability practices.

Culture & Benefits

  • Remote work supporting a global platform.
  • US-based employees are typically offered healthcare benefits, a 401(k) plan, short- and long-term disability benefits, and life and AD&D insurance.
  • Hands-on work with production SaaS systems serving thousands of customers across multiple regions and clouds.
  • Operational work includes continuous improvement of reliability, security, performance, and cost efficiency.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →