Назад
Company hidden
4 дня назад

Site Reliability Engineer Engineer

Формат работы
remote (Global)
Тип работы
fulltime
Грейд
senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer Engineer (AWS/Python): Building reliable, observable backend systems and SRE practices across AWS infrastructure and Python services with an accent on incident response, scalability, and operational resilience. Focus on designing observability, defining SLOs and error budgets, leading 24/7 incident response, and supporting AI workloads at scale.

Location: Remote only; daily overlap with client business hours is expected.

Company

hirify.global is a digital consulting partner that designs, builds, and modernizes digital products, platforms, and AI-powered experiences for client organizations.

What you will do

  • Design observability systems covering metrics, logging, tracing, alerting, SLOs, and error budgets.
  • Own 24/7 P0 on-call rotations with a 10-minute acknowledgment SLA and validate AI-generated incident reports.
  • Lead incident response, troubleshooting, root-cause analysis, postmortems, and reliability improvements.
  • Optimize capacity, latency, throughput, resource efficiency, and backend service performance.
  • Build self-healing automation, reduce operational toil, and create runbooks and architecture documentation.
  • Collaborate with developers, platform, operations, security, and client leadership while mentoring additional SRE engineers.

Requirements

  • 6+ years of hands-on experience in SRE, DevOps, or platform engineering at scale.
  • Deep AWS experience with ALB, ECS/Fargate, RDS Aurora, Lambda, and IAM.
  • Production Python backend experience, including debugging and optimizing containerized and Lambda services.
  • Experience building SRE programs, operating on-call rotations, conducting postmortems, and defining SLOs.
  • Strong background in observability, distributed systems, Kubernetes, Docker, infrastructure as code, and CI/CD.
  • Availability for client working-session overlap, reliable high-speed internet, and clear communication with technical stakeholders.

Nice to have

  • Experience with internal developer platforms, chaos engineering, resilience testing, or failure planning.
  • Cloud security, compliance, audit readiness, and FinOps experience.
  • AI infrastructure, LLM serving, agentic system operations, or AI-generated incident report workflows.
  • Experience mentoring engineers or leading technical design discussions.

Culture & Benefits

  • Fully remote collaboration with client and Modus teams.
  • Work with modern cloud platforms, Kubernetes, infrastructure as code, and observability tooling.
  • Continuous improvement through incident learning and operational excellence.
  • Direct impact on uptime, performance, reliability, and operational efficiency.
  • Opportunities to shape enterprise SRE practices and support AI services reliably at scale.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →