Назад
Company hidden
14 часов назад

Principal Site Reliability Engineer

7 000 - 12 000$
Формат работы
remote
Тип работы
fulltime
Грейд
senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Principal Site Reliability Engineer (SRE/AWS/Kubernetes): Architecting, upgrading, and building scalable infrastructure for production systems with an accent on reliability, recoverability, scalability, and distributed systems patterns. Focus on observability and alerting excellence, capacity planning and stress testing, and leading AI enablement to improve infrastructure reliability and developer productivity.

Location: Latin America

Salary: $7000 - $12000 USD per month (gross)

Company

hirify.global builds interest-free installment plans and modern shopping experiences for consumers and merchants.

What you will do

  • Architect, upgrade, design, and build scalable infrastructure using Kubernetes, AWS, and RDS (MySQL/Postgres).
  • Lead the infrastructure roadmap to improve reliability, recoverability, and scalability.
  • Drive capacity planning, benchmarking, and stress testing to identify bottlenecks and support business growth.
  • Define and enforce SLAs and alerts; improve anomaly detection and flexible alerting.
  • Lead AI enablement efforts to apply AI and automation for infrastructure reliability and developer productivity.
  • Establish engineering best practices for observability, security, and CI/CD; mentor engineers and translate business goals into technical roadmaps.

Requirements

  • 12+ years of professional software/infrastructure engineering experience, including significant SRE and backend experience.
  • Deployed significant changes to a production application or infrastructure configuration in the past 30 days.
  • Strong proficiency in Golang and experience building and maintaining RESTful APIs.
  • Expertise with SQL-based RDBMS (MySQL/PostgreSQL), including performance optimization at scale.
  • Proficiency in observability tools (Prometheus, Grafana, Datadog, New Relic).
  • Solid understanding of distributed systems design patterns (e.g., transactional outbox, event-driven architecture, stream processing, queues).

Culture & Benefits

  • High autonomy and authority to identify and resolve infrastructure and operational problems.
  • High standards for operational excellence, with a focus on fixing issues so they stay fixed.
  • Emphasis on calculated risk-taking and continuous improvement.
  • Mentorship and a culture of learning, innovation, and reliability.

Hiring process

  • Interviews focused on SRE/infrastructure leadership, production impact, and technical decision-making.
  • Discussion of experience with observability, alerting, distributed systems, and scaling high-traffic services.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →