2 дня назад
Staff Network Reliability Engineer, Cloud Operations (Kubernetes)
155 000 - 165 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Network Reliability Engineer, Cloud Operations (Kubernetes): Operating and improving Skylo's hybrid cloud infrastructure supporting a live commercial satellite connectivity network with an accent on Kubernetes, observability, storage, databases, and GitOps. Focus on defining SLOs, leading infrastructure incident response and RCA, automating remediation, and maintaining reliability across public and private cloud environments.
Location: Remote, United States
Salary: $155,000–$165,000 base salary per year plus equity
Company
operates a standards-based, cloud-native satellite connectivity platform that connects smartphones and IoT devices directly to satellites across consumer, automotive, and industrial IoT markets.
What you will do
- Own 24x7 cloud infrastructure health across GCP and on-premise private cloud environments, including GKE, bare-metal Kubernetes, storage, networking, and multi-cluster federation.
- Operate and improve the observability stack covering Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, Loki or ELK, Cloud Monitoring, and alert routing.
- Execute infrastructure runbooks and lead L3 incident response for Kubernetes, storage, database, networking, Pub/Sub, ArgoCD, and certificate-related failures.
- Maintain PostgreSQL replication, backup, restore, failover, and performance operations, as well as Redis cluster reliability.
- Define SLOs, track error budgets, lead capacity planning and RCA, and connect infrastructure reliability to network SLA commitments.
- Partner with Network Implementation and Platform Engineering on GitOps changes, operational readiness, automation, security hardening, and infrastructure roadmap requirements.
Requirements
- 8–10+ years of infrastructure engineering, SRE, or cloud operations experience in production 24x7 environments with direct Kubernetes on-call ownership.
- Deep Kubernetes expertise, including multi-cluster operations, node pools, RBAC, network policies, PVCs, CSI drivers, operators, and cluster upgrades.
- Hands-on experience with both public cloud, preferably GCP or AWS, and on-premise or private cloud infrastructure.
- Production experience with Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, PostgreSQL replication, Redis, ArgoCD or Flux CD, Helm, and Terraform or Ansible.
- Strong SRE fundamentals covering SLO/SLI/SLA definition, error budgets, toil reduction, capacity planning, incident response, and on-call operations.
- Ability to write diagnostic runbooks, deliver RCA documentation, and communicate structured infrastructure escalations.
Nice to have
- Telecom, 5G Core, vRAN, NTN, or satellite ground segment infrastructure experience.
- Production-scale Ceph or Rook, KubeVirt, Harvester, OpenStack, BGP, VXLAN, EVPN, or software-defined networking experience.
- Go or Python development for infrastructure automation, Kubernetes operators, or lifecycle integrations.
- FinOps experience or relevant certifications such as CKA, CKS, AWS Solutions Architect Professional, or RHCA.
Culture & Benefits
- Stock option-based equity program and competitive compensation.
- Medical, dental, vision, and retirement benefits.
- Monthly wellness and education reimbursements.
- Generous paid time off, holidays, and an opportunity to temporarily work abroad.
- Opportunity to operate a live commercial direct-to-device satellite network with an international engineering organization.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Backend Engineer (Infrastructure)
150 000 - 225 000$
CrowdStrike
1 день назад
Sr. Platform Engineer - Kubernetes (Remote)
140 000 - 215 000$
4 дня назад
Observability Engineer (Cloud)
100 000 - 160 000$
7 дней назад
Site Reliability Engineer (Kubernetes)
123 000 - 150 000$
5 дней назад
Senior DevOps Engineer (AWS)
140 000 - 175 000$
1 день назад
Senior Manager of Site Reliability Engineering (AWS)
155 000 - 170 000$