6 дней назад
Lead Site Reliability Engineer (FinTech)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer (FinTech): Defining, building, and operating always-on, low-latency, highly secure payment platforms with an accent on distributed systems, cloud platforms, observability, and reliability architecture. Focus on designing enterprise-grade observability and automation platforms, coordinating high-severity incident response, and driving resilience across regulated financial systems.
Location: Bangalore, India
Company
provides large-scale payment and financial technology platforms for mission-critical transactions.
What you will do
- Own reliability outcomes for real-time, distributed payment and transaction-processing platforms with strict SLAs, SLOs, and regulatory requirements.
- Define reliability architecture and standards across services, platforms, and infrastructure.
- Design and evolve observability platforms covering metrics, logs, traces, SLIs, and SLOs.
- Lead high-severity production incident response, root-cause analysis, and long-term systemic remediation.
- Drive SRE practices including error budgets, capacity modeling, resilience testing, graceful degradation, and operational readiness.
- Architect automation and self-service platforms, influence cloud migration and disaster recovery strategy, and mentor senior engineers and technical leads.
Requirements
- Deep software engineering experience building and operating large-scale, distributed, API-driven production systems.
- Expertise in observability, alerting, and reliability engineering with tools such as Prometheus, Grafana, Datadog, Splunk, or ELK.
- Strong knowledge of AWS, Azure, or GCP, infrastructure as code, platform automation, and cloud-native design patterns.
- Experience operating mission-critical systems in payments, fintech, banking, or similarly regulated environments.
- Hands-on experience with Linux/RHEL, Windows, Oracle RDBMS, and complex enterprise stacks.
- Staff-level scope, incident-management leadership, and the ability to align multiple teams and stakeholders.
Nice to have
- Automation and scripting with Python, Bash, Ansible, or similar tools.
- Experience building CI/CD platforms and release automation for high-risk production environments.
- Ownership of reliability strategy or platform initiatives across multiple teams or business units.
- Experience modernizing legacy financial systems into resilient, compliant cloud-native or hybrid architectures.
Culture & Benefits
- Work on high-impact payment platforms operating at massive scale.
- Define reliability strategy for systems where availability, correctness, and security are critical.
- Engineering culture focused on technical leadership, automation, excellence, and continuous learning.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Staff Site Reliability Engineer (AWS)
Akamai
4 дня назад
Senior Site Reliability Engineer (DevOps)
7 дней назад
Staff Site Reliability Engineer (Linux/Network Troubleshooting/Scripting) (AI)
6 дней назад
Incident Management Analyst
3 дня назад
Senior Cloud Operations Engineer, Infrastructure
102 400 - 153 200$