Назад
Company hidden
4 дня назад

Site Reliability Engineer Engineer

Формат работы
remote (Global)
Тип работы
fulltime
Грейд
senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer Engineer (AWS/Python/AI): Building reliable, observable backend systems across AWS infrastructure and Python services with an accent on incident response, SLOs, observability, and performance optimization. Focus on designing self-healing automation, leading 24/7 P0 incident operations, troubleshooting distributed systems, and supporting AI applications at scale.

Location: Remote; daily overlap with client business hours is expected.

Company

hirify.global is a global digital consulting partner that designs, builds, and modernizes digital products, platforms, and AI-powered experiences.

What you will do

  • Design observability systems covering metrics, logging, tracing, alerting, SLOs, and error budgets.
  • Own 24/7 P0 on-call rotations with a 10-minute acknowledgment SLA, including incident response, escalation, root-cause analysis, and postmortems.
  • Improve reliability, capacity, latency, throughput, performance, and resource efficiency across AWS infrastructure and Python services.
  • Build self-healing automation, reduce operational toil, and establish reliability standards for backend services.
  • Collaborate with developers, platform, operations, security, and client teams on architecture, delivery, and operational excellence.
  • Mentor and onboard SRE engineers as the team grows from 4 to 8–12 members.

Requirements

  • 6+ years of hands-on experience in SRE, DevOps, or Platform Engineering at scale.
  • Deep AWS experience with ALB, ECS/Fargate, RDS Aurora, Lambda, and IAM.
  • Production Python backend experience, including debugging and optimization of containerized services and Lambdas.
  • Experience building SRE programs, incident response processes, on-call rotations, postmortems, SLOs, observability, and distributed-system troubleshooting.
  • Experience with Kubernetes, Docker, infrastructure-as-code, CI/CD, deployment strategies, Git, and Linux/Unix administration.
  • Clear communication skills, documentation ability, and availability for regular client workshops and working sessions.

Nice to have

  • Experience with internal developer platforms, chaos engineering, resilience testing, or failure scenario planning.
  • Cloud security, compliance, audit readiness, and security hardening experience.
  • Experience with AI infrastructure, LLM serving, agentic system operations, or AI-generated incident report validation.
  • Cloud FinOps, cost optimization, Python performance tuning, or fast-scaling SRE program experience.

Culture & Benefits

  • Fully remote collaboration with global teams and client organizations.
  • Work on modern cloud platforms, Kubernetes, infrastructure-as-code, observability, and AI services.
  • Continuous improvement through incident learning, postmortems, documentation, and internal knowledge sharing.
  • Opportunities to influence enterprise reliability practices and operational standards.
  • Reliable high-speed internet is required for remote work.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →