Назад
Company hidden
2 месяца назад

Senior Lead Site Reliability Engineer (Fintech)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
India
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Lead Site Reliability Engineer (Fintech): Building and operating always-on, low-latency, highly secure payment platforms with an accent on distributed systems, cloud platforms, observability, and reliability engineering. Focus on designing resilient architectures, automating CI/CD and self-service operations, and resolving complex production issues across mission-critical financial systems.

Location: Pune, India

Company

hirify.global provides technology and payment platforms for financial services and large-scale financial transactions.

What you will do

  • Design and implement reliability architecture and engineering standards for scalable, resilient services, platforms, and infrastructure.
  • Build and evolve observability platforms covering metrics, logs, traces, SLOs, and SLIs.
  • Improve SRE practices including error budgets, capacity modeling, resilience testing, graceful degradation, and operational readiness.
  • Develop automation and self-service platforms that reduce operational toil and enable safe production releases.
  • Design and optimize secure CI/CD pipelines, deployment automation, quality gates, and rollback strategies.
  • Troubleshoot complex production issues, perform root-cause analysis, and deliver long-term reliability improvements with engineering and infrastructure teams.

Requirements

  • 10–15 years of IT experience.
  • Strong software engineering experience designing, building, and operating large-scale distributed, API-driven production systems.
  • Hands-on CI/CD and release engineering experience with Azure DevOps, GitHub Actions, Jenkins, Harness, or equivalent tools.
  • Expertise in observability, alerting, and reliability engineering using Prometheus, Grafana, Datadog, Splunk, ELK, or equivalent ecosystems.
  • Experience with AWS, Azure, or GCP, infrastructure as code, platform automation, and cloud-native design patterns.
  • Hands-on experience with Linux, Windows, databases, enterprise technology stacks, and highly available mission-critical platforms.

Nice to have

  • Automation and scripting with Python, Bash, Ansible, or PowerShell; C#/.NET experience is a plus.
  • Ownership of reliability strategy or platform initiatives spanning multiple teams or business units.
  • Experience modernizing legacy financial systems into cloud-native or hybrid architectures with a focus on resilience and compliance.

Culture & Benefits

  • Modern international work environment with a collaborative and motivated team.
  • Broad professional education and personal development opportunities.
  • Competitive salary and benefits.
  • Career development tools, resources, and opportunities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →