4 дня назад
Site Reliability Engineer Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer Engineer (AWS/Python/AI): Building reliable, observable backend systems across AWS infrastructure and Python services with an accent on incident response, SLOs, observability, and performance optimization. Focus on designing self-healing automation, leading 24/7 P0 incident operations, troubleshooting distributed systems, and supporting AI applications at scale.
Location: Remote; daily overlap with client business hours is expected.
Company
is a global digital consulting partner that designs, builds, and modernizes digital products, platforms, and AI-powered experiences.
What you will do
- Design observability systems covering metrics, logging, tracing, alerting, SLOs, and error budgets.
- Own 24/7 P0 on-call rotations with a 10-minute acknowledgment SLA, including incident response, escalation, root-cause analysis, and postmortems.
- Improve reliability, capacity, latency, throughput, performance, and resource efficiency across AWS infrastructure and Python services.
- Build self-healing automation, reduce operational toil, and establish reliability standards for backend services.
- Collaborate with developers, platform, operations, security, and client teams on architecture, delivery, and operational excellence.
- Mentor and onboard SRE engineers as the team grows from 4 to 8–12 members.
Requirements
- 6+ years of hands-on experience in SRE, DevOps, or Platform Engineering at scale.
- Deep AWS experience with ALB, ECS/Fargate, RDS Aurora, Lambda, and IAM.
- Production Python backend experience, including debugging and optimization of containerized services and Lambdas.
- Experience building SRE programs, incident response processes, on-call rotations, postmortems, SLOs, observability, and distributed-system troubleshooting.
- Experience with Kubernetes, Docker, infrastructure-as-code, CI/CD, deployment strategies, Git, and Linux/Unix administration.
- Clear communication skills, documentation ability, and availability for regular client workshops and working sessions.
Nice to have
- Experience with internal developer platforms, chaos engineering, resilience testing, or failure scenario planning.
- Cloud security, compliance, audit readiness, and security hardening experience.
- Experience with AI infrastructure, LLM serving, agentic system operations, or AI-generated incident report validation.
- Cloud FinOps, cost optimization, Python performance tuning, or fast-scaling SRE program experience.
Culture & Benefits
- Fully remote collaboration with global teams and client organizations.
- Work on modern cloud platforms, Kubernetes, infrastructure-as-code, observability, and AI services.
- Continuous improvement through incident learning, postmortems, documentation, and internal knowledge sharing.
- Opportunities to influence enterprise reliability practices and operational standards.
- Reliable high-speed internet is required for remote work.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (AWS)
6 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
6 дней назад
Site Reliability Engineer
2 750 - 4 800€
10 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
5 дней назад
Senior Site Reliability Engineer (AWS)
130 000 - 155 000$
5 дней назад