Site Reliability Engineering Manager (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Remote in the United States, except Delaware, Nevada, Ohio, Oregon, Hawaii, New Mexico, and West Virginia. Roles may be based internationally in some cases; employment contracts outside the US and Ireland are powered by Rippling.
Salary: $175,000–$238,000 annual base pay, with a typical midpoint of $207,000. Total compensation also includes equity and benefits.
Company
provides logistics technology and infrastructure that connect merchants with carriers through a single API and dashboard.
What you will do
- Lead and develop a platform-focused SRE team through technical mentorship, career development, and performance management.
- Build internal platforms, Kubernetes infrastructure, deployment tooling, and self-service capabilities for product engineering teams.
- Manage observability platforms covering metrics, logs, traces, dashboards, and reliability measurement.
- Establish SLO, SLI, and error-budget frameworks while improving deployment success, build times, and developer experience.
- Drive automation, infrastructure cost optimization, capacity planning, and operational excellence across the cloud platform.
- Lead Sev1 incident response, manage the on-call rotation, and design disaster recovery, security, and compliance capabilities.
Requirements
- 3+ years of engineering management experience and 9+ years as a software or systems engineer.
- Expertise building internal platforms and tooling for other engineering teams, including production Kubernetes platforms.
- Deep experience with AWS or GCP, including networking, compute, storage, and managed services.
- Experience with CI/CD and deployment automation, infrastructure as code, and observability platforms.
- Proficiency in Python, Go, or a similar programming language for tooling and automation.
- Experience with reliability frameworks, disaster recovery, infrastructure security, compliance, and cross-functional communication.
Nice to have
- Experience with GitHub Actions, GitLab CI, ArgoCD, Flux, Terraform, Pulumi, CloudFormation, Prometheus, Grafana, ELK, Datadog, or New Relic.
Culture & Benefits
- Remote-first, globally distributed work environment with flexible working hours.
- Medical, dental, and vision coverage, with 90% covered by the company including dependents.
- Flexible vacation policy, a company-wide winter slowdown, and three volunteer days off.
- Work-from-home stipend, pet coverage, charity donation matching, and individual learning support.
- Professional development programs, coaching, and regular company and local gatherings.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →