Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
DevOps Engineer (Observability) (OpenTelemetry/AWS/Kubernetes): Rebuilding Twilio’s observability platform with an accent on scalable telemetry collection, real-time analytics, and unified logs, metrics, traces, and profiling. Focus on designing high-scale platform components, improving incident response and performance analysis, and balancing reliability, cost, and usability across distributed systems.
Location: Remote from Ireland; occasional travel may be required for project or team in-person meetings.
Company
Twilio provides communications solutions that help businesses and developers create personalized customer experiences.
What you will do
- Lead the architecture and delivery of observability platform components focused on reliability, scalability, and usability.
- Build consistent workflows for logs, metrics, traces, and continuous profiling.
- Design developer-friendly tooling and APIs for incident response, performance analysis, and platform debugging.
- Drive the observability transformation using centralized S3-based data lakes, OpenTelemetry instrumentation, and ClickHouse-backed query engines.
- Collaborate with product teams, SREs, and developer experience groups to integrate observability into engineering workflows.
- Provide architectural guidance, mentor engineers, and align cross-team efforts with long-term platform goals.
Requirements
- Proven experience building and scaling observability systems such as logging platforms, metrics pipelines, tracing infrastructure, or profiling tools.
- Experience designing high-scale telemetry systems and handling high-cardinality data and telemetry correlation.
- Proficiency in at least one modern programming language, such as Go, Python, or Java.
- Strong understanding of distributed systems and observability challenges in microservice-based environments.
- Experience with AWS, Kubernetes, and infrastructure-as-code tools.
- Ability to make forward-looking technical decisions, provide architectural guidance, and lead through ambiguity.
Nice to have
- Familiarity with ClickHouse, Grafana Mimir, Athena, or equivalent log and metrics querying systems.
- Contributions to open-source observability tools or communities.
- Experience building cost visibility or FinOps tooling for cloud compute and telemetry pipelines.
Culture & Benefits
- Remote-first work with opportunities for team gatherings, functional off-sites, and customer meetings.
- Competitive pay and generous paid time off.
- Parental and wellness leave.
- Healthcare and a retirement savings program.
- Support for volunteering and community donations.
Hiring process
- Hiring decisions are made by Twilio employees, with AI used to support process efficiency.
- Formal interviews are part of the recruitment process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Staff Observability Engineer (AI)
290 000 - 375 000PLN
10 дней назад
DevOps & SRE Engineer (Kubernetes)
100 000 - 150 000$
13 дней назад
Member of Technical Staff (Observability & Reliability)
14 дней назад
Manager, Site Reliability Engineering (Cloud Infrastructure)
12 дней назад
Principal Site Reliability Engineer, Platform Engineering
223 200 - 380 400$
3 часа назад