Назад
Company hidden
2 дня назад

Senior Site Reliability Engineer (Fintech)

129 000 - 175 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Fintech): Building and operating reliable, scalable mobile and digital platforms with an accent on observability, automation, incident management, and distributed systems. Focus on designing Python and shell tooling, improving CI/CD and monitoring, implementing SLOs and error budgets, and reducing operational toil across high-availability environments.

Location: On site in Southlake, Texas, United States. Applicants must be authorized to work in the United States full time without employer sponsorship.

Salary: USD $129,000–$175,000 per year, plus eligibility for bonus or incentive opportunities.

Company

hirify.global is a financial services company focused on transforming the finance industry through digital and mobile platforms.

What you will do

  • Respond to production alerts and incident escalations, leading triage, resolution, root cause analysis, and post-incident reviews.
  • Participate in an on-call rotation supporting high-availability systems.
  • Define monitoring, telemetry, dashboards, alerting, and observability practices across systems.
  • Design Python and shell scripting automation to improve resilience, service recovery, certificate management, and system maintenance.
  • Improve CI/CD pipelines and embed reliability practices into the software development lifecycle.
  • Influence engineering teams, establish operational guardrails, and mentor junior engineers in SRE practices.

Requirements

  • 10+ years of software development and site reliability engineering experience with cloud-native architectures and distributed systems.
  • 8+ years of DevOps and/or SRE experience focused on production operations, automation, and large-scale reliability.
  • 8+ years of experience with CI/CD pipelines, observability, and monitoring or telemetry platforms.
  • 5+ years implementing and scaling SRE practices, including SLOs, monitoring strategies, incident reviews, and automation improvements.
  • Experience supporting production-grade, high-availability distributed systems and building reliability tooling.
  • Bachelor of Science in Computer Science or a related field, or equivalent work experience.

Nice to have

  • Python or Java programming experience for scalable services and APIs.
  • Application performance monitoring experience, preferably with Splunk.
  • Kubernetes, Terraform, and public cloud experience with AWS, GCP, or Azure.
  • Knowledge of cloud infrastructure components such as compute, storage, networking, load balancing, DNS, and security architectures.
  • Experience leading cross-functional technical strategy and communicating complex concepts to varied audiences.

Culture & Benefits

  • Regular in-person collaboration and an on-site working arrangement for this role.
  • 401(k) with company match and an employee stock purchase plan.
  • Paid vacation, volunteering time, and a 28-day sabbatical after five years for eligible positions.
  • Paid parental leave, adoption and family-building benefits.
  • Health, dental, and vision insurance, plus tuition reimbursement.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →