Назад
Company hidden
6 дней назад

Site Reliability Engineer (Kubernetes)

Формат работы
remote (Global)
Тип работы
fulltime
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Kubernetes): Building and operating cloud infrastructure for backend services and a carrier-as-a-service telecom platform with an accent on automation, observability, and high availability. Focus on designing deployment and recovery systems, handling production incidents, and optimizing distributed services under heavy load.

Location: Remote

Company

hirify.global is building a carrier-as-a-service telecom platform that enables clients to launch customizable mobile networks, connect users and devices across carriers and technologies, and extract value from telecom data.

What you will do

  • Design and implement cloud platforms supporting backend services.
  • Automate deployments, scaling, recovery, and other technical operations.
  • Monitor and maintain mission-critical production infrastructure for high availability.
  • Participate in on-call rotations, incident response, and blameless postmortems.
  • Provide tools and operational enablement for Engineering, Telecom, and Data Engineering teams.

Requirements

  • Strong understanding of Linux/Unix systems, including processes, filesystems, memory management, and networking.
  • Programming and scripting experience with Python, Go, Ruby, Bash, or Perl.
  • Experience with infrastructure provisioning, containers, orchestration, and cloud platforms, including Terraform, Ansible, Docker, Kubernetes, AWS, Google Cloud, or Azure.
  • Experience with monitoring, alerting, log analysis, dashboards, incident management, CI/CD pipelines, and on-call operations.
  • Knowledge of TCP/IP, DNS, HTTP/HTTPS, load balancing, firewalls, deployment strategies, high availability, failover, IAM, and zero-trust principles.
  • Experience with distributed systems, databases, virtualization, configuration management, load testing, and performance optimization.

Nice to have

  • Experience with Prometheus, Grafana, Datadog, Kafka, Cassandra, Elasticsearch, Jaeger, OpenTelemetry, ELK, Splunk, VMware, KVM, or SaltStack.
  • Experience building custom monitoring tools and complex automation scripts.

Culture & Benefits

  • Remote full-time work within an engineering department.
  • Participation in a culture of continuous improvement and blameless postmortems.
  • Collaboration with Engineering, Telecom, and Data Engineering teams.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →