21 час назад
Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes/Cloud): Improving the reliability and resilience of Semrush infrastructure and applications with an accent on failure recovery, observability, SLOs, and scalable platform architecture. Focus on designing full-stack platform solutions, building Go/Python automation tooling, debugging production systems, and leading critical incidents.
Company
is a brand visibility platform that combines SEO authority and AI-driven insights to help marketers improve their online presence.
What you will do
- Lead changes to common engineering practices and improve the reliability of critical systems.
- Induce application failures, recover services, and lead critical incident response.
- Debug applications with metrics and implement traces and additional observability.
- Establish and refine SLOs, cost dashboards, and security hardening initiatives with stakeholders.
- Design scalable, reliable full-stack platform solutions from concept through production.
- Build Go/Python tooling to automate operations, mentor engineers, and participate in candidate interviews.
Requirements
- 3+ years of experience as a Site Reliability Engineer.
- Experience with Kubernetes and cloud providers.
- Professional engineering experience with Python or Go.
- Strong understanding of application failures, recovery, metrics-based debugging, traces, and observability.
- Willingness to participate in on-call rotations, typically one week every 2–3 weeks, including possible overnight incidents.
- Strong communication skills and a collaborative approach.
Nice to have
- Knowledge of GCP.
Culture & Benefits
- Collaborative work across engineering, data, product, and design teams.
- Unlimited paid time off.
- Hobby and team-building budget.
- Employee Support Program.
- Financial aid following the loss of a family member.
- Employee Resource Groups and an inclusive workplace commitment.
Hiring process
- Applications are reviewed by the Talent Acquisition team, with feedback targeted within three working days.
- Interviews include a detailed discussion of professional background, working style, and personality.
- Online interviews are expected to be conducted from a laptop or desktop with the camera enabled.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Site Reliability Engineer (AWS)
1 день назад
Senior Site Reliability Engineer (Kubernetes)
Affirm
6 дней назад
Senior Site Reliability Engineer (SRE & Platform Reliability)
308 000 - 428 000PLN
2 дня назад
DevOps / SRE Engineer
3 дня назад
Sr. Site Reliability Engineer
125 000 - 145 000$
21 час назад