Назад
Company hidden
обновлено 18 дней назад

Service Reliability Engineer

Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
Colombia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Service Reliability Engineer (SRE/Cloud): Ensuring the reliability, availability, scalability, and performance of mission-critical airline platforms across cloud and on-premise environments with an accent on automation, observability, incident management, and distributed systems. Focus on building monitoring and recovery solutions, troubleshooting microservices and APIs, improving deployment strategies, and conducting root cause analysis in follow-the-sun production operations.

Location: Hybrid work at the Bogota office, Colombia; participation in on-call and follow-the-sun rotations is required.

Company

hirify.global is a global travel technology company developing platforms and systems that support airline operations and serve travelers worldwide.

What you will do

  • Ensure high availability, scalability, resilience, and performance of mission-critical production systems.
  • Manage incidents, problems, and changes using ITIL practices; conduct root cause analysis and drive issue resolution.
  • Automate operational tasks, deployments, and recovery processes, including blue/green and canary deployment strategies.
  • Build and improve monitoring, alerting, and observability across metrics, logs, and traces while tracking SLOs and SLAs.
  • Operate cloud, on-premise, containerized, and distributed environments across Linux and/or Windows platforms.
  • Troubleshoot microservices, APIs, databases, and multi-stack applications while documenting playbooks and mentoring junior engineers.

Requirements

  • Experience as a Site Reliability Engineer or similar production engineering role; typically 3+ years, adaptable to seniority.
  • Strong production operations, incident management, reliability engineering, and mission-critical systems experience.
  • Experience with Linux and/or Windows Server, cloud platforms such as Azure, AWS, or GCP, and distributed systems.
  • Knowledge of observability tools such as Grafana, Prometheus, ELK, or Splunk, plus CI/CD and automation tooling.
  • Experience with Kubernetes or OpenShift, scripting or programming in Python, Shell, Go, or similar, and Ansible or comparable automation tools.
  • Good communication skills in English and the ability to collaborate across globally distributed teams under production pressure.

Nice to have

  • .NET and C# application debugging, runtime configuration, and IIS administration.
  • Experience with Kafka, caching, messaging systems, and SQL or NoSQL databases.

Culture & Benefits

  • Competitive remuneration with individual and company annual bonuses.
  • Paid vacation and holidays, health insurance, equity, and other benefits.
  • Professional development through online technical and soft-skills learning hubs.
  • Diverse and inclusive workplace with a global culture and opportunities to support travel technology used by millions of travelers.

Hiring process

  • Create a candidate profile, upload an English CV or resume, and apply through the official application process.
  • The application process takes no longer than 10 minutes.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →