Назад
Company hidden
8 дней назад

Senior Site Reliability Engineer

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Philippines
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Cloud/SRE): Operating and improving reliable global unified communications infrastructure with an accent on incident response, automation, observability, and security. Focus on designing reliability tooling, driving SLO-based decisions, reducing operational toil, and leading complex production responses.

Location: Manila, Philippines

Company

hirify.global provides unified communications and customer experience solutions that connect customers and teams globally.

What you will do

  • Own platform reliability across global unified communications infrastructure and lead incident response for the assigned subsystem.
  • Triage complex production issues, perform scheduled maintenance, and redesign operational processes to prevent recurring failures.
  • Lead blameless post-mortems, track corrective actions, and convert recurring incidents into engineering improvements and actionable bug reports.
  • Design automation and tooling, reduce manual toil, address technical debt, and deliver infrastructure initiatives through two-week sprint cycles.
  • Define and track SLIs, SLOs, and SLAs; build and maintain Grafana and OCI Log Analytics dashboards and improve alert quality.
  • Provide technical leadership and mentorship, run workshops, document operational knowledge, and support projects involving one or two other engineers.

Requirements

  • 6+ years of experience in site reliability, platform operations, or infrastructure engineering, including operating production systems at scale.
  • Advanced Linux systems administration skills, including distributed services, log analysis, systemctl, and network diagnostics.
  • Hands-on experience with at least one major cloud provider: OCI, AWS, GCP, or Azure.
  • Strong on-call and incident response experience, including structured triage, stakeholder communication, and post-mortem follow-through.
  • Ability to script in Python or Bash and strong knowledge of SRE concepts, including SLIs, SLOs, error budgets, and toil measurement.
  • Demonstrated technical leadership, end-to-end feature ownership, mentoring ability, and an AI-forward approach to daily work.

Nice to have

  • Experience with Oracle Cloud Infrastructure, including compute, networking, Log Analytics, and Object Storage.
  • Familiarity with VoIP and SIP infrastructure, including registration, trunking, and call signaling.
  • Knowledge of Prometheus, Grafana, PagerDuty, OCI Log Analytics, and Ansible.
  • Experience with large-scale infrastructure migrations in multi-tenant SaaS environments.

Culture & Benefits

  • Participation in a shared on-call rotation of approximately one week per month.
  • Escalation is encouraged, with an expectation to coordinate incident response rather than work alone.
  • Collaboration with Support, Sales, Sales Engineering, NOC, Professional Services, and Engineering teams.
  • Equal employment opportunities and reasonable accommodation for employees and applicants with disabilities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →