Назад
Company hidden
22 часа назад

Senior Site Reliability Engineer

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Site Reliability Engineer (SRE/Observability): Improving the reliability of high-load frontend and backend services through incident response, monitoring, and service-level objectives with an accent on distributed systems, observability, and operational automation. Focus on investigating complex production incidents, reducing detection and recovery times, and maintaining availability across critical services.

Location: Tbilisi; hybrid work format with flexible working hours.

Company

hirify.global is a financial analysis platform providing charting, market data, collaboration, and publishing tools to more than 100 million users worldwide.

What you will do

  • Investigate production incidents, perform root cause analysis, and drive corrective actions through resolution.
  • Develop and improve monitoring, alerting, observability, and service health measurement.
  • Define and maintain SLI/SLOs, availability requirements, error budgets, and SLA compliance for assigned services.
  • Analyze service performance, resource utilization, capacity risks, and reliability gaps.
  • Create runbooks, troubleshooting guides, recovery procedures, and operational documentation.
  • Automate diagnostics and incident response, participate in on-call rotations, and train duty engineers.

Requirements

  • Professional experience as an SRE, Reliability Engineer, Production Engineer, Operations Engineer, or in a similar role.
  • Hands-on experience with production incident investigation, root cause analysis, and postmortem processes.
  • Strong understanding of monitoring, alerting, observability, metrics, logs, and distributed tracing.
  • Knowledge of SLA, SLI, SLO, and error budget concepts.
  • Understanding of distributed systems and high-load environments.
  • Experience automating operational and repetitive tasks and maintaining runbooks.

Nice to have

  • CKA or CKAD certification.
  • Experience with Prometheus, Grafana, OpenTelemetry, or similar observability platforms.
  • Experience implementing reliability, incident management, and problem management practices.
  • Understanding of the Google Four Golden Signals.
  • Experience using AI-assisted tools for incident investigation and operating large-scale distributed systems.

Culture & Benefits

  • Flexible working hours and a hybrid work environment.
  • Well-equipped offices for focused and collaborative work.
  • Relocation support and private health insurance.
  • Learning, mentorship, and long-term career development.
  • Performance-based bonuses, hirify.global Premium access, and regular team events.
  • Global distributed collaboration across a diverse international workforce.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →