4 дня назад
Site Reliability Engineer Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer Engineer (AWS/Python): Building reliable, observable backend systems and SRE practices across AWS infrastructure and Python services with an accent on incident response, scalability, and operational resilience. Focus on designing observability, defining SLOs and error budgets, leading 24/7 incident response, and supporting AI workloads at scale.
Location: Remote only; daily overlap with client business hours is expected.
Company
is a digital consulting partner that designs, builds, and modernizes digital products, platforms, and AI-powered experiences for client organizations.
What you will do
- Design observability systems covering metrics, logging, tracing, alerting, SLOs, and error budgets.
- Own 24/7 P0 on-call rotations with a 10-minute acknowledgment SLA and validate AI-generated incident reports.
- Lead incident response, troubleshooting, root-cause analysis, postmortems, and reliability improvements.
- Optimize capacity, latency, throughput, resource efficiency, and backend service performance.
- Build self-healing automation, reduce operational toil, and create runbooks and architecture documentation.
- Collaborate with developers, platform, operations, security, and client leadership while mentoring additional SRE engineers.
Requirements
- 6+ years of hands-on experience in SRE, DevOps, or platform engineering at scale.
- Deep AWS experience with ALB, ECS/Fargate, RDS Aurora, Lambda, and IAM.
- Production Python backend experience, including debugging and optimizing containerized and Lambda services.
- Experience building SRE programs, operating on-call rotations, conducting postmortems, and defining SLOs.
- Strong background in observability, distributed systems, Kubernetes, Docker, infrastructure as code, and CI/CD.
- Availability for client working-session overlap, reliable high-speed internet, and clear communication with technical stakeholders.
Nice to have
- Experience with internal developer platforms, chaos engineering, resilience testing, or failure planning.
- Cloud security, compliance, audit readiness, and FinOps experience.
- AI infrastructure, LLM serving, agentic system operations, or AI-generated incident report workflows.
- Experience mentoring engineers or leading technical design discussions.
Culture & Benefits
- Fully remote collaboration with client and Modus teams.
- Work with modern cloud platforms, Kubernetes, infrastructure as code, and observability tooling.
- Continuous improvement through incident learning and operational excellence.
- Direct impact on uptime, performance, reliability, and operational efficiency.
- Opportunities to shape enterprise SRE practices and support AI services reliably at scale.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
8 дней назад
Site Reliability Engineer, Tech Lead (AI)
5 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
8 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
8 дней назад
Senior DevOps / Site Reliability Engineer (SRE) (Cybersecurity)
165 000 - 215 000$
7 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$