22 часа назад
Site Reliability Engineer (Observability)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Observability): Operating and improving a shared telemetry platform for metrics, logs, traces, alerting, dashboards, and profiling with an accent on distributed systems, automation, and production reliability. Focus on troubleshooting data pipelines and capacity issues, managing infrastructure as code and containerized workloads, and improving incident response through reusable automation and runbooks.
Location: LATAM, Argentina, Brazil, Canada, Chile, Colombia, Paraguay, Peru, or Uruguay
Company
is a global crypto platform offering spot trading, margin, futures, staking, and OTC services for individual and institutional clients.
What you will do
- Operate and improve the shared telemetry platform for metrics, logs, traces, alerting, dashboards, and profiling.
- Maintain Prometheus-compatible monitoring, VictoriaMetrics, Grafana, Vector, Splunk, Loki, Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope systems.
- Deploy and manage telemetry services across multiple environments using Terraform, Terragrunt, and container orchestration.
- Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, availability, latency, and capacity issues.
- Build reusable configuration and automation for dashboards, alerts, and telemetry integrations.
- Participate in incident response and on-call rotations, write runbooks, and improve operations based on incident learnings.
Requirements
- 3+ years of experience as a Site Reliability Engineer, Platform or Infrastructure Engineer, Observability Engineer, or similar production engineer.
- Experience managing production systems at scale that collect, process, store, and serve telemetry.
- Experience with Prometheus or a compatible monitoring stack, including metrics collection, querying, and alerting.
- Experience troubleshooting distributed production systems and working with Terraform and CI/CD.
- Experience operating containerized workloads with Nomad, Kubernetes, or similar platforms.
- Strong scripting or programming, incident response, documentation, collaboration skills, and comfort using AI tools and agents.
Nice to have
- Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, OpenTelemetry, PromQL, or LogQL.
- Experience with dashboards or alerts as code and Kubernetes operators or CRDs.
- Experience with Consul, Vault, AWS, on-premises infrastructure, high-volume logging, streaming, or data pipelines.
- Background in regulated or financial services environments with strict change management and audit trails.
Culture & Benefits
- Work with software engineers, platform teams, security engineers, and SREs on mission-critical services.
- Applications are accepted on an ongoing basis unless a specific deadline is stated.
- Job-related skills or work-style assessments may be included in the hiring process.
- Hiring is based on merit, with an emphasis on diverse backgrounds and perspectives.
Hiring process
- Complete job-related skills or work-style assessments if requested.
- Assessment results are considered alongside experience and interviews.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
18 часов назад
Staff Production Engineer (AI)
209 000 - 253 000$
Windsurf
7 дней назад
Site Reliability Engineer (AI)
Helsing
6 дней назад
Site Reliability Engineer (AI)
19 часов назад
Site Reliability Engineer - Vice President
3 дня назад
Site Reliability Engineer (AWS)
3 дня назад