3 дня назад
Site Reliability Engineer (AWS)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AWS/SaaS): Building and operating resilient, highly available infrastructure for Redwood’s mission-critical SaaS platform with an accent on AWS networking, EKS/Kubernetes, Infrastructure as Code, and observability. Focus on leading incident response, automating remediation, designing multi-region cloud environments, and improving reliability through SLOs, error budgets, and root cause analysis.
Location: United States (Remote)
Company
Software provides a SaaS-first orchestration platform for automating mission-critical business and IT processes across ERP, hybrid cloud, data, and agentic AI systems.
What you will do
- Manage alerts, monitor system health, and participate in a shared 24x7 on-call rotation for critical SaaS incidents.
- Lead incident response, root cause analysis, and blameless post-mortems, defining corrective actions to prevent recurrence.
- Design, deploy, and maintain multi-region, multi-account AWS infrastructure using Terraform, Helm, EKS/Kubernetes, Docker, and Docker Swarm.
- Design and troubleshoot AWS networking, including VPCs, subnetting, routing, security groups, load balancers, VPN, and Transit Gateway connectivity.
- Develop automation, self-healing checks, CI/CD pipelines, monitoring, alerting, and observability across metrics, logs, and traces.
- Collaborate with engineering, Support, Customer Success, Migration, and Professional Services teams to minimize customer impact during changes.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, or cloud infrastructure roles supporting production SaaS platforms.
- Strong AWS engineering and networking experience, including IAM, VPC design, Route 53, ALB/NLB, security groups/NACLs, Transit Gateway, and VPN connectivity.
- Proficiency with EKS/Kubernetes, Terraform, Helm, Docker, and Docker Swarm.
- Hands-on experience with Datadog, including APM, infrastructure monitoring, synthetic testing, and dashboards.
- Linux administration and Bash and/or Python scripting skills, plus production experience with relational databases such as AWS RDS/PostgreSQL.
- Experience with DevSecOps, high-availability SaaS support, observability, and communicating technical issues to technical and customer-facing audiences.
Nice to have
- Experience with Grafana and Prometheus.
Culture & Benefits
- Work remotely from the United States.
- Participate in a team-shared 24x7 on-call rotation for critical incidents.
- Work within a customer-focused environment centered on reliability, collaboration, curiosity, and ownership.
- Use blameless post-mortems and error budgets to balance development velocity with platform stability.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Site Reliability Engineer (AWS)
120 000 - 185 000$
9 дней назад
Senior Site Reliability Engineer (Kubernetes/AWS)
4 дня назад
Sr. Site Reliability Engineer
160 000 - 180 000$
10 дней назад
Staff Site Reliability Engineer (Kubernetes)
3 часа назад
DevOps/SRE Engineer
3 дня назад