Назад
Company hidden
2 месяца назад

Monitoring and Observability Engineer

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Armenia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Monitoring and Observability Engineer (Prometheus/Kubernetes): Designing and operating end-to-end observability infrastructure for metrics, logs, traces, and synthetic monitoring across on-premise and cloud environments with an accent on custom agents, distributed tracing, and actionable alerting. Focus on scaling monitoring platforms, reducing alert noise, improving MTTR, and solving complex visibility challenges across multi-layered infrastructure.

Location: Yerevan, Armenia; hybrid work environment

Company

hirify.global is seeking an engineer to shape observability infrastructure supporting platform reliability and performance.

What you will do

  • Design and architect observability solutions for metrics, logs, traces, and synthetic monitoring across on-premise and cloud environments.
  • Build and maintain custom monitoring agents and exporters for internal services and infrastructure.
  • Develop dashboards, alerts, and runbooks that provide actionable insights for operations and engineering teams.
  • Own deployment, scaling, performance tuning, and capacity planning for the monitoring stack.
  • Implement distributed tracing pipelines and define observability standards across teams.
  • Support incident response, root-cause analysis, tooling evaluation, and collaboration with platform, networking, and application teams.

Requirements

  • 5+ years of experience in monitoring, observability, or SRE roles with hands-on ownership of observability infrastructure.
  • Deep expertise in at least two of Prometheus or VictoriaMetrics, Elastic Stack, Grafana, Loki, Jaeger, Tempo, or OpenTelemetry.
  • Production-grade programming experience with Go, Python, or similar, including custom monitoring agents or exporters.
  • Strong Linux administration, networking, TCP/IP, DNS, VLANs, routing, traffic analysis, firewall, kernel, and process-management knowledge.
  • Experience operating monitoring systems in Kubernetes and using IaC or GitOps tools such as Terraform, Ansible, Flux, or ArgoCD.
  • Strong incident response, scripting, automation, log aggregation, and alerting strategy experience.

Nice to have

  • Experience with distributed tracing, service mesh observability, APM, synthetic monitoring, or time-series database query optimization.
  • Familiarity with Kafka-based telemetry pipelines and AWS, GCP, Azure, multi-cloud, or hybrid monitoring.
  • Background in capacity planning and performance engineering.

Culture & Benefits

  • Open, feedback-driven culture supporting continuous professional growth.
  • Distributed, hybrid work environment with teams across multiple time zones.
  • Competitive compensation and well-being benefits.
  • Opportunity to shape observability strategy and pilot modern tooling and practices.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →