13 дней назад
Member of Technical Staff (Observability & Reliability)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff (Observability & Reliability) (OpenTelemetry/Kubernetes): Evolving observability and reliability systems across cloud and customer-hosted dataplanes with an accent on telemetry, SLOs, incident response, and deployment health. Focus on detecting state drift, monitoring ephemeral ML workloads, reducing MTTR, and controlling telemetry costs.
Location: Remote in São Paulo, Brazil
Company
operates cloud and customer-hosted dataplanes that support real-time operational decisions and batch inference workloads.
What you will do
- Evolve the observability stack for logs, metrics, traces, and alerting.
- Ensure cloud and on-premise dataplanes report releases, health, heartbeats, telemetry, and usage to the control plane.
- Bring telemetry into customer Kubernetes clusters through outbound-only agent connections.
- Detect desired-state drift and monitor deployment and runtime agents.
- Provide visibility into ephemeral workloads, including Ray clusters used for batch inference.
- Define SLOs, lead incident response and postmortems, reduce MTTR, and optimize telemetry costs.
Requirements
- Deep experience with OpenTelemetry and observability backends.
- Hands-on experience with SLOs, error budgets, actionable alerting, and incident management.
- Strong experience with Kubernetes and infrastructure as code using Terraform and Helm.
- Experience operating software in environments that are not fully controlled.
- Ability to write and review production-quality code and operate the systems built.
Nice to have
- Experience shipping software to customer-hosted Kubernetes environments, including Helm and outbound-only connectivity.
- Experience with GCP/GKE or AWS/EKS.
- Production experience with multi-node or multi-cluster ML workloads.
- Experience in financial services or regulated environments.
Culture & Benefits
- Remote work arrangement in São Paulo.
- Member of Technical Staff model with ownership of systems and outcomes.
- Success measured through serving availability, incident reduction, MTTR, state consistency, and active dataplane agents.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
employ.city
4 дня назад
Senior DevOps/SRE (Observability) Engineer
Valletta.Software | AI-Care
12 часов назад
Senior DevOps / SRE Support Engineer (LATAM)
5 000 - 5 500$
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
13 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
14 дней назад
Senior Staff Engineer, Site Reliability Engineering
3 часа назад