Назад
Company hidden
24 часа назад

Site Reliability Engineer (Prometheus/Grafana)

Формат работы
remote (только Europe)
Тип работы
fulltime
Английский
b2
Страна
Ukraine/Poland/Romania +1 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Prometheus/Grafana): Keeping a global social gaming platform stable by monitoring infrastructure and applications, diagnosing incidents, and optimizing observability with an accent on metrics, logs, alerts, and large-scale production systems. Focus on performing root cause analysis, reducing alert noise, coordinating incident response, and maintaining reliability across shift-based European coverage.

Location: Remote from Ukraine, Bulgaria, Poland, or Romania; shift-based schedule with European time alignment.

Company

hirify.global is a fast-scaling product development company building social gaming technology for millions of players worldwide.

What you will do

  • Monitor infrastructure and application health across a multi-site production environment using Prometheus and Grafana.
  • Analyze metrics, logs, and alerts to identify and resolve issues before they affect players.
  • Perform root cause analysis for incidents and document solutions for future reference.
  • Fine-tune monitoring systems, improve observability, and reduce alert noise.
  • Collaborate with development and operations teams during incident response and reviews.
  • Contribute to continuous improvement in a 24×7 SRE environment.

Requirements

  • Technical foundation in system administration, DevOps, or technical support.
  • Strong understanding of server, network, and application performance metrics, including CPU, memory, latency, and RPS.
  • Experience with log analysis tools such as Elasticsearch, Kibana, Loki, or Splunk.
  • Hands-on experience configuring Prometheus, Alertmanager, or Datadog.
  • Proficiency with Linux systems and command-line operations.
  • Willingness to work shift-based schedules during non-business hours aligned with European time.

Nice to have

  • Kubernetes experience.

Culture & Benefits

  • Fully remote work arrangement within the specified countries.
  • Paid vacation and sick leave.
  • Company events and a vibrant work culture.
  • Opportunities for growth in a fast-scaling business.
  • Exposure to modern monitoring technologies and large-scale systems.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →