7 дней назад
Senior Cloud Infrastructure and Networking
125 000 - 135 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Cloud Infrastructure and Networking (Kubernetes/GCP): Operating and improving Skylo's hybrid cloud infrastructure for a live satellite connectivity network with an accent on Kubernetes, observability, storage, databases, and GitOps. Focus on resolving complex production incidents, defining SLOs and error budgets, automating recovery, and maintaining reliable infrastructure across public and private clouds.
Location: Remote, US
Base salary: $125,000–$135,000 per year
Company
provides standards-based satellite connectivity that connects smartphones and IoT devices directly to satellites through a cloud-native commercial NTN vRAN platform.
What you will do
- Own 24x7 infrastructure health across GCP and on-premise private cloud environments, including GKE, bare-metal Kubernetes, storage, networking, and compute platforms.
- Operate the observability stack built with Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, Loki or ELK, and Pub/Sub-based alerting.
- Maintain PostgreSQL and Redis reliability, including replication, backup and restore, failover testing, performance monitoring, and persistence configuration.
- Serve as the L3 escalation authority for Cloud Infrastructure incidents and lead troubleshooting bridges through resolution or structured engineering handoff.
- Define SLOs, SLIs, error budgets, capacity plans, runbooks, SOPs, and root-cause analyses for infrastructure components.
- Partner with Network Implementation, Platform Engineering, security, Core NRE, and RAN NRE teams on GitOps changes, operational readiness, automation, and infrastructure hardening.
Requirements
- 5+ years of infrastructure engineering, SRE, or cloud operations experience in a production 24x7 environment with direct Kubernetes on-call ownership.
- Deep Kubernetes expertise covering multi-cluster operations, node pools, RBAC, network policies, PVCs, CSI drivers, operators, and production upgrades.
- Hands-on experience with public cloud and private cloud infrastructure, including GCP or AWS, bare-metal Kubernetes, KVM, or hyperconverged platforms.
- Production experience with Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, alerting pipelines, PostgreSQL replication, Redis, ArgoCD or Flux, Helm, Terraform, or Ansible.
- Strong knowledge of SRE practices, Linux and container internals, networking, incident response, runbook authorship, RCA documentation, and operational communication.
- Must be based in the United States; the role is remote and participates in a global 24x7 on-call rotation.
Nice to have
- Experience with telecom, 5G Core, vRAN, NTN, or satellite ground-segment infrastructure.
- Production-scale Ceph or Rook, KubeVirt, Harvester, OpenStack, BGP, VXLAN, EVPN, or software-defined networking experience.
- Strong Go or Python development experience for automation tooling, Kubernetes operators, or infrastructure integrations.
- FinOps experience or certifications such as CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect.
Culture & Benefits
- Flexible remote work with employees across three continents.
- Medical, dental, vision, retirement, paid time off, holidays, and stock option-based equity.
- Monthly wellness and education reimbursement allowances.
- Opportunity to work on a live direct-to-device satellite network and collaborate with specialists across software, hardware, telecom, satellite, and network virtualization.
- Inclusive culture with a focus on transparency and diversity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Replit
13 дней назад
Site Reliability Engineer
210 000 - 275 000$
Prenosis
7 дней назад
Senior Software/SRE Engineer (Go)
120 000 - 150 000$
13 дней назад
Site Reliability Engineer (AWS/Kubernetes)
90 000 - 120 000GBP
Valletta.Software | AI-Care
12 часов назад
Senior DevOps / SRE Support Engineer (LATAM)
5 000 - 5 500$
Replit
13 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
9 дней назад