14 часов назад
Senior Site Reliability Engineer (GCP)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (GCP): Building and operating scalable, reliable infrastructure for a cloud-based supply chain platform with an accent on GKE, Cloud Run, AlloyDB, Terraform, and observability. Focus on automating infrastructure and operational workflows, improving reliability and cost efficiency, and designing incident response and disaster-recovery practices for distributed systems.
Location: Remote, United States
Employment type: Full time
Company
provides a cloud-based supply chain and commerce platform covering order management, warehouse management, transportation, fulfillment, and consumer experience.
What you will do
- Architect and implement scalable infrastructure on Google Cloud Platform across GKE, Cloud Run, AlloyDB, networking, and IAM.
- Own Terraform infrastructure-as-code modules, organization policies, and reusable platform patterns.
- Manage Kubernetes workloads, including performance tuning, capacity planning, resource optimization, and cost reduction.
- Build monitoring, alerting, and observability with Datadog, and define reliability signals for owned services.
- Design disaster-recovery and business-continuity strategies and validate their effectiveness.
- Develop GitHub Actions CI/CD pipelines, automate operational workflows, support production incidents, and improve reliability practices.
Requirements
- 5+ years of experience in SRE, platform, or infrastructure engineering.
- Strong production experience with GCP, including GKE, Cloud Run, AlloyDB, networking, and IAM.
- Advanced Docker, Kubernetes, and Terraform experience, including workload scaling and reusable modules.
- Productivity in TypeScript, Python, Go, or a similar programming language for tooling and automation.
- Experience with actionable observability, distributed-systems fundamentals, Git workflows, incident management, and post-mortems.
- Ability to communicate technical trade-offs, collaborate across teams, take ownership, and use AI-assisted development tools responsibly.
Nice to have
- PostgreSQL operations, database migrations and scaling, Redis, ClickHouse, Kafka, Redpanda, or Pub/Sub experience.
- Cloud-cost optimization, GCP certifications, Cloudflare Workers, or multi-cloud and hybrid-architecture experience.
Culture & Benefits
- Small, fast-moving SRE team with broad ownership and a short path from decisions to production.
- Shared ownership, clear communication, technical design reviews, and collaboration with development and data teams.
- Participation in on-call support for critical systems and post-incident improvement work.
- Opportunity to shape reliability, automation, and infrastructure practices during rapid company growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
17 часов назад
Senior Site Reliability Engineer (Kubernetes)
12 часов назад
Site Reliability Engineer Staff (Cloud Infrastructure)
3 дня назад
Senior SRE (AI)
6 дней назад
Senior Site Reliability Engineer (GovCloud)
117 000 - 209 330$
5 дней назад
Senior Site Reliability Engineer (Kubernetes)
93 700 - 138 700$
6 дней назад
Site Reliability Engineer
87 400 - 123 400$