8 дней назад
Senior Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Cloud/SRE): Operating and improving reliable global unified communications infrastructure with an accent on incident response, automation, observability, and security. Focus on designing reliability tooling, driving SLO-based decisions, reducing operational toil, and leading complex production responses.
Location: Manila, Philippines
Company
provides unified communications and customer experience solutions that connect customers and teams globally.
What you will do
- Own platform reliability across global unified communications infrastructure and lead incident response for the assigned subsystem.
- Triage complex production issues, perform scheduled maintenance, and redesign operational processes to prevent recurring failures.
- Lead blameless post-mortems, track corrective actions, and convert recurring incidents into engineering improvements and actionable bug reports.
- Design automation and tooling, reduce manual toil, address technical debt, and deliver infrastructure initiatives through two-week sprint cycles.
- Define and track SLIs, SLOs, and SLAs; build and maintain Grafana and OCI Log Analytics dashboards and improve alert quality.
- Provide technical leadership and mentorship, run workshops, document operational knowledge, and support projects involving one or two other engineers.
Requirements
- 6+ years of experience in site reliability, platform operations, or infrastructure engineering, including operating production systems at scale.
- Advanced Linux systems administration skills, including distributed services, log analysis, systemctl, and network diagnostics.
- Hands-on experience with at least one major cloud provider: OCI, AWS, GCP, or Azure.
- Strong on-call and incident response experience, including structured triage, stakeholder communication, and post-mortem follow-through.
- Ability to script in Python or Bash and strong knowledge of SRE concepts, including SLIs, SLOs, error budgets, and toil measurement.
- Demonstrated technical leadership, end-to-end feature ownership, mentoring ability, and an AI-forward approach to daily work.
Nice to have
- Experience with Oracle Cloud Infrastructure, including compute, networking, Log Analytics, and Object Storage.
- Familiarity with VoIP and SIP infrastructure, including registration, trunking, and call signaling.
- Knowledge of Prometheus, Grafana, PagerDuty, OCI Log Analytics, and Ansible.
- Experience with large-scale infrastructure migrations in multi-tenant SaaS environments.
Culture & Benefits
- Participation in a shared on-call rotation of approximately one week per month.
- Escalation is encouraged, with an expectation to coordinate incident response rather than work alone.
- Collaboration with Support, Sales, Sales Engineering, NOC, Professional Services, and Engineering teams.
- Equal employment opportunities and reasonable accommodation for employees and applicants with disabilities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Senior Site Reliability Engineer (SRE)
7 дней назад
Site Reliability Engineer (Cloud Infrastructure)
7 дней назад
Site Reliability Engineer - Warehousing IT Operations (Cloud Infrastructure)
6 дней назад
Senior Site Reliability Engineer (FinTech)
23 часа назад
Platform Site Reliability Engineer (SRE)
5 дней назад