5 дней назад
Systems Engineer - Level III (AI/AIOps)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Systems Engineer - Level III (AI/AIOps): Maintaining the availability, reliability, and operational health of business-critical infrastructure, applications, and services in a hybrid technology environment with an accent on monitoring, incident response, and operational automation. Focus on reducing alert noise, coordinating major incident response, improving runbooks and escalation workflows, and adopting AI-enabled troubleshooting and remediation practices.
Location: Manila, Philippines
Company
promotes collaboration, expertise, and innovation while providing business services through Arch Global Services (Philippines) Inc.
What you will do
- Respond rapidly to infrastructure, application, network, and security-related operational events.
- Monitor, triage, prioritize, communicate, and escalate global incidents within agreed service levels.
- Use monitoring, observability, alerting, and incident management platforms to detect service degradation and reduce alert noise.
- Manage incident tickets with accurate impact details, timelines, actions, communications, and resolution documentation.
- Lead or support major incident bridges with technical teams, service owners, and business stakeholders.
- Improve runbooks, escalation paths, alert quality, response procedures, and AI-enabled operational practices.
Requirements
- Experience managing multiple operational priorities, incidents, and project-related tasks in a fast-paced environment.
- Strong troubleshooting, analytical thinking, problem-solving, communication, and stakeholder coordination skills.
- Working knowledge of Windows Server, Linux, networking, DNS, DHCP, firewalls, VPN, load balancing, and cloud or hybrid environments.
- Familiarity with monitoring, observability, ticketing, and incident response tools such as SolarWinds, ServiceNow, PagerDuty, Splunk, Dynatrace, Azure Monitor, or similar platforms.
- Understanding of alert correlation, event enrichment, escalation workflows, runbooks, operational automation, and AI-enabled operations.
- Bachelor's degree or equivalent experience; willingness to participate in rotating shifts, off-hours support, and on-call duties.
Nice to have
- Experience with virtualization, cloud platforms, and enterprise infrastructure support.
- Exposure to PowerShell, Python, REST APIs, workflow automation, or low-code automation platforms.
- Experience improving alert tuning, runbooks, knowledge bases, and post-incident reviews.
Culture & Benefits
- Collaborative environment focused on expertise, innovation, and continuous improvement.
- Opportunity to work with observability, automation, AIOps, predictive monitoring, and generative AI capabilities.
- Role includes participation in rotating shifts, off-hours support, and on-call coverage.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →