Senior Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Site Reliability Engineer (SRE/Observability): Improving the reliability of high-load frontend and backend services through incident response, monitoring, and service-level objectives with an accent on distributed systems, observability, and operational automation. Focus on investigating complex production incidents, reducing detection and recovery times, and maintaining availability across critical services.
Location: Tbilisi; hybrid work format with flexible working hours.
Company
is a financial analysis platform providing charting, market data, collaboration, and publishing tools to more than 100 million users worldwide.
What you will do
- Investigate production incidents, perform root cause analysis, and drive corrective actions through resolution.
- Develop and improve monitoring, alerting, observability, and service health measurement.
- Define and maintain SLI/SLOs, availability requirements, error budgets, and SLA compliance for assigned services.
- Analyze service performance, resource utilization, capacity risks, and reliability gaps.
- Create runbooks, troubleshooting guides, recovery procedures, and operational documentation.
- Automate diagnostics and incident response, participate in on-call rotations, and train duty engineers.
Requirements
- Professional experience as an SRE, Reliability Engineer, Production Engineer, Operations Engineer, or in a similar role.
- Hands-on experience with production incident investigation, root cause analysis, and postmortem processes.
- Strong understanding of monitoring, alerting, observability, metrics, logs, and distributed tracing.
- Knowledge of SLA, SLI, SLO, and error budget concepts.
- Understanding of distributed systems and high-load environments.
- Experience automating operational and repetitive tasks and maintaining runbooks.
Nice to have
- CKA or CKAD certification.
- Experience with Prometheus, Grafana, OpenTelemetry, or similar observability platforms.
- Experience implementing reliability, incident management, and problem management practices.
- Understanding of the Google Four Golden Signals.
- Experience using AI-assisted tools for incident investigation and operating large-scale distributed systems.
Culture & Benefits
- Flexible working hours and a hybrid work environment.
- Well-equipped offices for focused and collaborative work.
- Relocation support and private health insurance.
- Learning, mentorship, and long-term career development.
- Performance-based bonuses, Premium access, and regular team events.
- Global distributed collaboration across a diverse international workforce.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →