9 дней назад
Site Reliability Engineer (AWS)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AWS): Operating and strengthening AWS production infrastructure with an accent on observability, infrastructure automation, incident response, and resilient systems. Focus on improving Kubernetes and EKS workloads, building Terraform-based infrastructure, maintaining CI/CD pipelines, and solving complex reliability and performance issues.
Location: Remote-first role based in the United States; attendance at multiple company-wide and team-specific onsites is expected throughout the year. An office in Washington, DC is available for optional use.
Company
provides a fundraising platform for nonprofit educational institutions, serving colleges, universities, and K-12 schools.
What you will do
- Operate, maintain, and improve production infrastructure in AWS.
- Build infrastructure as code with Terraform and support Kubernetes and Amazon EKS workloads.
- Improve observability through dashboards, alerts, service-level indicators, and platforms such as New Relic.
- Investigate incidents, identify root causes, implement durable fixes, and participate in the shared 24/7 on-call rotation.
- Partner with product engineers on application resilience, performance, reliability, and production-readiness initiatives.
- Maintain CI/CD pipelines, automate operational work, and create runbooks, system diagrams, and troubleshooting documentation.
Requirements
- Approximately 5+ years of experience in software engineering, infrastructure, systems engineering, SRE, platform engineering, DevOps, or equivalent practical experience.
- Hands-on experience operating production workloads in AWS and maintaining infrastructure with Terraform or a similar tool.
- Experience with New Relic, Datadog, or another modern observability platform, as well as production incident troubleshooting and on-call rotations.
- Experience building or maintaining CI/CD pipelines and making targeted changes to application or automation code.
- Working knowledge of Linux, networking, distributed systems, and relational databases.
- Strong communication skills and the ability to explain root causes, manage scoped projects, and communicate risks and tradeoffs.
Nice to have
- Experience with Ruby or Ruby on Rails.
- PostgreSQL administration or performance tuning.
- Kubernetes, Amazon EKS, Redis, OpenSearch, or Amazon RDS experience.
- Experience operating enterprise SaaS products at scale or working with payments, fintech, regulated systems, SOC 2, or similar compliance programs.
- Familiarity with SLOs, SLIs, error budgets, capacity modeling, or load testing.
Culture & Benefits
- Remote-first work with a distributed team across more than 30 U.S. states.
- Regular company-wide and team-specific onsite gatherings, partner institution visits, and retreats.
- Blameless postmortems and a collaborative approach to reliability work.
- Inclusive, supportive, and learning-oriented work environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Site Reliability Engineering Manager (AWS/Kubernetes)
205 000 - 255 000$
9 дней назад
Cloud Site Reliability Engineer (AWS)
120 000 - 130 000$
Replit
13 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
1 день назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
Replit
13 дней назад
Site Reliability Engineer
210 000 - 275 000$
11 дней назад
Senior/Lead Site Reliability Engineer (Federal)
159 000 - 230 000$