Назад
Company hidden
2 дня назад

Staff Network Reliability Engineer, Cloud Operations (Kubernetes)

155 000 - 165 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Network Reliability Engineer, Cloud Operations (Kubernetes): Operating and improving Skylo's hybrid cloud infrastructure supporting a live commercial satellite connectivity network with an accent on Kubernetes, observability, storage, databases, and GitOps. Focus on defining SLOs, leading infrastructure incident response and RCA, automating remediation, and maintaining reliability across public and private cloud environments.

Location: Remote, United States

Salary: $155,000–$165,000 base salary per year plus equity

Company

hirify.global operates a standards-based, cloud-native satellite connectivity platform that connects smartphones and IoT devices directly to satellites across consumer, automotive, and industrial IoT markets.

What you will do

  • Own 24x7 cloud infrastructure health across GCP and on-premise private cloud environments, including GKE, bare-metal Kubernetes, storage, networking, and multi-cluster federation.
  • Operate and improve the observability stack covering Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, Loki or ELK, Cloud Monitoring, and alert routing.
  • Execute infrastructure runbooks and lead L3 incident response for Kubernetes, storage, database, networking, Pub/Sub, ArgoCD, and certificate-related failures.
  • Maintain PostgreSQL replication, backup, restore, failover, and performance operations, as well as Redis cluster reliability.
  • Define SLOs, track error budgets, lead capacity planning and RCA, and connect infrastructure reliability to network SLA commitments.
  • Partner with Network Implementation and Platform Engineering on GitOps changes, operational readiness, automation, security hardening, and infrastructure roadmap requirements.

Requirements

  • 8–10+ years of infrastructure engineering, SRE, or cloud operations experience in production 24x7 environments with direct Kubernetes on-call ownership.
  • Deep Kubernetes expertise, including multi-cluster operations, node pools, RBAC, network policies, PVCs, CSI drivers, operators, and cluster upgrades.
  • Hands-on experience with both public cloud, preferably GCP or AWS, and on-premise or private cloud infrastructure.
  • Production experience with Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, PostgreSQL replication, Redis, ArgoCD or Flux CD, Helm, and Terraform or Ansible.
  • Strong SRE fundamentals covering SLO/SLI/SLA definition, error budgets, toil reduction, capacity planning, incident response, and on-call operations.
  • Ability to write diagnostic runbooks, deliver RCA documentation, and communicate structured infrastructure escalations.

Nice to have

  • Telecom, 5G Core, vRAN, NTN, or satellite ground segment infrastructure experience.
  • Production-scale Ceph or Rook, KubeVirt, Harvester, OpenStack, BGP, VXLAN, EVPN, or software-defined networking experience.
  • Go or Python development for infrastructure automation, Kubernetes operators, or lifecycle integrations.
  • FinOps experience or relevant certifications such as CKA, CKS, AWS Solutions Architect Professional, or RHCA.

Culture & Benefits

  • Stock option-based equity program and competitive compensation.
  • Medical, dental, vision, and retirement benefits.
  • Monthly wellness and education reimbursements.
  • Generous paid time off, holidays, and an opportunity to temporarily work abroad.
  • Opportunity to operate a live commercial direct-to-device satellite network with an international engineering organization.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →