Назад
Company hidden
13 дней назад

Member of Technical Staff (Observability & Reliability)

Формат работы
remote (только Brazil)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Brazil
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff (Observability & Reliability) (OpenTelemetry/Kubernetes): Evolving observability and reliability systems across cloud and customer-hosted dataplanes with an accent on telemetry, SLOs, incident response, and deployment health. Focus on detecting state drift, monitoring ephemeral ML workloads, reducing MTTR, and controlling telemetry costs.

Location: Remote in São Paulo, Brazil

Company

hirify.global operates cloud and customer-hosted dataplanes that support real-time operational decisions and batch inference workloads.

What you will do

  • Evolve the observability stack for logs, metrics, traces, and alerting.
  • Ensure cloud and on-premise dataplanes report releases, health, heartbeats, telemetry, and usage to the control plane.
  • Bring telemetry into customer Kubernetes clusters through outbound-only agent connections.
  • Detect desired-state drift and monitor deployment and runtime agents.
  • Provide visibility into ephemeral workloads, including Ray clusters used for batch inference.
  • Define SLOs, lead incident response and postmortems, reduce MTTR, and optimize telemetry costs.

Requirements

  • Deep experience with OpenTelemetry and observability backends.
  • Hands-on experience with SLOs, error budgets, actionable alerting, and incident management.
  • Strong experience with Kubernetes and infrastructure as code using Terraform and Helm.
  • Experience operating software in environments that are not fully controlled.
  • Ability to write and review production-quality code and operate the systems built.

Nice to have

  • Experience shipping software to customer-hosted Kubernetes environments, including Helm and outbound-only connectivity.
  • Experience with GCP/GKE or AWS/EKS.
  • Production experience with multi-node or multi-cluster ML workloads.
  • Experience in financial services or regulated environments.

Culture & Benefits

  • Remote work arrangement in São Paulo.
  • Member of Technical Staff model with ownership of systems and outcomes.
  • Success measured through serving availability, incident reduction, MTTR, state consistency, and active dataplane agents.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →