3 дня назад
Lead Site Reliability Engineer (Azure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer (Azure/FinTech): Leading the development of reliable, scalable, and high-performing cloud infrastructure and SRE solutions for investment management products with an accent on Azure, observability, automation, and incident management. Focus on defining SLOs and error budgets, implementing anomaly detection and self-healing mechanisms, managing production incidents, and mentoring SRE engineers.
Location: Warsaw, Poland; hybrid model with 2 days in the office and 3 days remote. Occasional remote work is available across Poland and internationally for limited periods under internal policy. Regular and evening shifts, weekend support, and on-call rotations may be required.
Company
International provider of integrated investment management solutions for investment and asset managers, operating within the Group.
What you will do
- Lead the development of SRE solutions covering monitoring, alerting, anomaly detection, self-healing, and reliability testing.
- Design reliability, scalability, performance, capacity planning, resource management, and automation strategies across products and onboarding pipelines.
- Build automation and tooling to reduce manual operational work, improve developer experience, and increase engineering velocity.
- Define and manage SLOs and error budgets with engineering teams while driving observability and operational excellence.
- Lead incident response, root cause analysis, disaster recovery, configuration management, and platform readiness activities.
- Mentor junior SREs, guide architectural decisions, collaborate with product teams and stakeholders, and contribute to on-call support.
Requirements
- 5–8+ years of experience in SRE or cloud infrastructure leadership, plus a Bachelor’s or Master’s degree in Computer Science or a related field.
- Extensive production-grade expertise in Microsoft Azure and cloud-native reliability practices.
- Experience with Infrastructure as Code using Bicep, ARM, and Terraform.
- Hands-on experience with Azure Monitor, Application Insights, DataDog, Log Analytics, OpenTelemetry, and distributed tracing.
- Experience with incident response, ITIL-based problem and change management, identity and access protocols such as SAML, OAuth, and OIDC, and synthetic monitoring with Playwright or an equivalent tool.
- Knowledge of Kubernetes, Docker, networking, APIs, PowerShell, Bash, Kusto, SQL, Cosmos DB, PostgreSQL, and AI/ML-based anomaly detection.
Nice to have
- Familiarity with SimCorp Dimension and Salesforce.
- Experience managing onboarding projects alongside live production operations.
Culture & Benefits
- Flexible working hours and a hybrid work model with a modern office near Wilanowska metro station.
- Base salary with an annual bonus structure; compensation ranges are disclosed during the initial candidate engagement.
- Employer-paid healthcare, travel insurance, group life insurance, and a Multisport card with employer contribution.
- Holiday allowance for a two-week vacation and access to professional training and courses.
- Career development in an international environment, language classes, integration events, volunteering initiatives, and employee-led clubs.
Hiring process
- Applications are reviewed continuously through the career site.
- The approximate CV review time is three weeks, followed by a candidate feedback process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →