Назад
Company hidden
6 дней назад

Site Reliability Engineering Manager (Kubernetes)

175 000 - 238 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineering Manager (Kubernetes): Building platforms, tooling, and infrastructure that enable product teams to operate reliable, performant, and scalable services with an accent on observability, deployment automation, infrastructure management, and developer self-service. Focus on leading SRE engineers, establishing SLO/SLI frameworks, optimizing cloud infrastructure, and ensuring disaster recovery, security, and operational resilience.

Location: Remote in the United States, except Delaware, Nevada, Ohio, Oregon, Hawaii, New Mexico, and West Virginia. Roles may be based internationally in some cases; employment contracts outside the US and Ireland are powered by Rippling.

Salary: $175,000–$238,000 annual base pay, with a typical midpoint of $207,000. Total compensation also includes equity and benefits.

Company

hirify.global provides logistics technology and infrastructure that connect merchants with carriers through a single API and dashboard.

What you will do

  • Lead and develop a platform-focused SRE team through technical mentorship, career development, and performance management.
  • Build internal platforms, Kubernetes infrastructure, deployment tooling, and self-service capabilities for product engineering teams.
  • Manage observability platforms covering metrics, logs, traces, dashboards, and reliability measurement.
  • Establish SLO, SLI, and error-budget frameworks while improving deployment success, build times, and developer experience.
  • Drive automation, infrastructure cost optimization, capacity planning, and operational excellence across the cloud platform.
  • Lead Sev1 incident response, manage the on-call rotation, and design disaster recovery, security, and compliance capabilities.

Requirements

  • 3+ years of engineering management experience and 9+ years as a software or systems engineer.
  • Expertise building internal platforms and tooling for other engineering teams, including production Kubernetes platforms.
  • Deep experience with AWS or GCP, including networking, compute, storage, and managed services.
  • Experience with CI/CD and deployment automation, infrastructure as code, and observability platforms.
  • Proficiency in Python, Go, or a similar programming language for tooling and automation.
  • Experience with reliability frameworks, disaster recovery, infrastructure security, compliance, and cross-functional communication.

Nice to have

  • Experience with GitHub Actions, GitLab CI, ArgoCD, Flux, Terraform, Pulumi, CloudFormation, Prometheus, Grafana, ELK, Datadog, or New Relic.

Culture & Benefits

  • Remote-first, globally distributed work environment with flexible working hours.
  • Medical, dental, and vision coverage, with 90% covered by the company including dependents.
  • Flexible vacation policy, a company-wide winter slowdown, and three volunteer days off.
  • Work-from-home stipend, pet coverage, charity donation matching, and individual learning support.
  • Professional development programs, coaching, and regular company and local gatherings.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →