Назад
Company hidden
14 часов назад

Senior Site Reliability Engineer (GCP)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Europe/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (GCP): Building and operating scalable, reliable infrastructure for a cloud-based supply chain platform with an accent on GKE, Cloud Run, AlloyDB, Terraform, and observability. Focus on automating infrastructure and operational workflows, improving reliability and cost efficiency, and designing incident response and disaster-recovery practices for distributed systems.

Location: Remote, United States

Employment type: Full time

Company

hirify.global provides a cloud-based supply chain and commerce platform covering order management, warehouse management, transportation, fulfillment, and consumer experience.

What you will do

  • Architect and implement scalable infrastructure on Google Cloud Platform across GKE, Cloud Run, AlloyDB, networking, and IAM.
  • Own Terraform infrastructure-as-code modules, organization policies, and reusable platform patterns.
  • Manage Kubernetes workloads, including performance tuning, capacity planning, resource optimization, and cost reduction.
  • Build monitoring, alerting, and observability with Datadog, and define reliability signals for owned services.
  • Design disaster-recovery and business-continuity strategies and validate their effectiveness.
  • Develop GitHub Actions CI/CD pipelines, automate operational workflows, support production incidents, and improve reliability practices.

Requirements

  • 5+ years of experience in SRE, platform, or infrastructure engineering.
  • Strong production experience with GCP, including GKE, Cloud Run, AlloyDB, networking, and IAM.
  • Advanced Docker, Kubernetes, and Terraform experience, including workload scaling and reusable modules.
  • Productivity in TypeScript, Python, Go, or a similar programming language for tooling and automation.
  • Experience with actionable observability, distributed-systems fundamentals, Git workflows, incident management, and post-mortems.
  • Ability to communicate technical trade-offs, collaborate across teams, take ownership, and use AI-assisted development tools responsibly.

Nice to have

  • PostgreSQL operations, database migrations and scaling, Redis, ClickHouse, Kafka, Redpanda, or Pub/Sub experience.
  • Cloud-cost optimization, GCP certifications, Cloudflare Workers, or multi-cloud and hybrid-architecture experience.

Culture & Benefits

  • Small, fast-moving SRE team with broad ownership and a short path from decisions to production.
  • Shared ownership, clear communication, technical design reviews, and collaboration with development and data teams.
  • Participation in on-call support for critical systems and post-incident improvement work.
  • Opportunity to shape reliability, automation, and infrastructure practices during rapid company growth.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →