Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Software Engineer (L4) (Site Reliability Engineering): Building and operating reliable, performant, and recoverable production services for a global communications platform with an accent on distributed systems, observability, and production resilience. Focus on designing failure-tolerant architectures, defining SLIs and SLOs, automating reliability improvements, and leading incident response across multiple engineering teams.
Location: Remote, based in Ireland; occasional travel may be required for project or team meetings.
Company
Twilio provides communications solutions that enable businesses and developers to create personalized customer experiences.
What you will do
- Own the reliability posture of production services, including availability, latency, capacity, performance, monitoring, and alerting.
- Define and operate against SLIs, SLOs, and error budgets to guide engineering priorities.
- Design failure-tolerant systems, validate recovery paths, and make production changes safer to deploy and roll back.
- Write, review, test, and deploy production code and automation that improves service reliability and reduces operational overhead.
- Lead incident response, post-mortem analysis, debugging, troubleshooting, and follow-up work to prevent recurrence.
- Drive cross-team reliability projects from conception to completion and provide technical leadership through design feedback, code review, and mentorship.
Requirements
- 8+ years of related engineering experience, including substantial experience in reliability, infrastructure, or platform engineering.
- Accountability for production systems, including on-call responsibilities and ownership during service failures.
- Strong software engineering fundamentals and experience shipping production code.
- Experience with SLIs, SLOs, error budgets, incident command, post-mortems, capacity planning, and observability.
- Experience driving changes across multiple teams and building alignment without formal authority.
- Experience with large-scale distributed systems in a cloud environment.
Nice to have
- Infrastructure-as-code, container orchestration, and GitOps-style delivery experience.
- Experience with multi-region architecture, failure-domain design, or regional expansion.
- Background in chaos engineering, game days, or proactive resilience validation.
Culture & Benefits
- Remote-first work with occasional in-person team, project, or customer meetings.
- Competitive pay and generous time off.
- Parental and wellness leave.
- Healthcare and a retirement savings program, with offerings varying by location.
- Opportunities to support volunteering and community initiatives.
Hiring process
- Hiring decisions are made by Twilio employees, with artificial intelligence used to support process efficiency.
- A formal interview process is required; offers are not made without interviews.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Anthropic
8 дней назад
Staff Software Engineer (AI Reliability Engineering)
325 000 - 390 000GBP
8 дней назад
Senior Site Reliability Engineer (SRE) – Infrastructure & Systems (Remote)
32 000PLN
7 дней назад
Senior Site Reliability Engineer (SRE) – Infrastructure & Systems (B2B)
12 дней назад
Staff Engineer (SRE)
95 800 - 185 000$
10 дней назад
Staff Engineer, Software Engineering (SRE Availability and Incident Management)
100 000 - 230 000$
8 дней назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$