Назад
Company hidden
3 дня назад

Site Reliability Engineer (AWS)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AWS/SaaS): Building and operating resilient, highly available infrastructure for Redwood’s mission-critical SaaS platform with an accent on AWS networking, EKS/Kubernetes, Infrastructure as Code, and observability. Focus on leading incident response, automating remediation, designing multi-region cloud environments, and improving reliability through SLOs, error budgets, and root cause analysis.

Location: United States (Remote)

Company

hirify.global Software provides a SaaS-first orchestration platform for automating mission-critical business and IT processes across ERP, hybrid cloud, data, and agentic AI systems.

What you will do

  • Manage alerts, monitor system health, and participate in a shared 24x7 on-call rotation for critical SaaS incidents.
  • Lead incident response, root cause analysis, and blameless post-mortems, defining corrective actions to prevent recurrence.
  • Design, deploy, and maintain multi-region, multi-account AWS infrastructure using Terraform, Helm, EKS/Kubernetes, Docker, and Docker Swarm.
  • Design and troubleshoot AWS networking, including VPCs, subnetting, routing, security groups, load balancers, VPN, and Transit Gateway connectivity.
  • Develop automation, self-healing checks, CI/CD pipelines, monitoring, alerting, and observability across metrics, logs, and traces.
  • Collaborate with engineering, Support, Customer Success, Migration, and Professional Services teams to minimize customer impact during changes.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, or cloud infrastructure roles supporting production SaaS platforms.
  • Strong AWS engineering and networking experience, including IAM, VPC design, Route 53, ALB/NLB, security groups/NACLs, Transit Gateway, and VPN connectivity.
  • Proficiency with EKS/Kubernetes, Terraform, Helm, Docker, and Docker Swarm.
  • Hands-on experience with Datadog, including APM, infrastructure monitoring, synthetic testing, and dashboards.
  • Linux administration and Bash and/or Python scripting skills, plus production experience with relational databases such as AWS RDS/PostgreSQL.
  • Experience with DevSecOps, high-availability SaaS support, observability, and communicating technical issues to technical and customer-facing audiences.

Nice to have

  • Experience with Grafana and Prometheus.

Culture & Benefits

  • Work remotely from the United States.
  • Participate in a team-shared 24x7 on-call rotation for critical incidents.
  • Work within a customer-focused environment centered on reliability, collaboration, curiosity, and ownership.
  • Use blameless post-mortems and error budgets to balance development velocity with platform stability.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →