9 дней назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI/Voice Communications): Building and operating reliable production infrastructure for cloud communications, Unified Communications, and AI-enabled services with an accent on observability, scalability, and real-time media quality. Focus on designing resilient Voice/UC and AI service paths, automating production-readiness checks, and solving complex failure, latency, dependency, and recovery challenges.
Location: United Kingdom; primarily remote with occasional visits to the Bristol or London office
Company
provides cloud communications and collaboration technology, including Voice/Unified Communications services and AI-powered capabilities.
What you will do
- Operate and improve production environments, monitoring availability, system health, capacity, performance, and failure modes.
- Build software, automation, and infrastructure systems for distributed cloud applications and services.
- Develop observability across infrastructure, applications, communications telemetry, and AI-service metrics.
- Improve the reliability of Voice/UC platforms, including signaling, media flows, call quality, latency, jitter, packet loss, failover, and end-to-end availability.
- Partner with development, Voice/UC, and AI engineering teams on testing, releases, production readiness, and AI-enabled communications capabilities.
- Design recovery procedures, graceful degradation, dependency isolation, retry and fallback patterns, and automated validation.
Requirements
- Bachelor's degree in computer science or another technical or scientific discipline, or equivalent practical experience.
- 4–7 years of experience in production operations, systems engineering, SRE/DevOps, CI/CD, software deployment, and production-system maintenance.
- Experience with Agile, DevOps, CI/CD pipelines, infrastructure automation, monitoring, observability, distributed systems, cloud infrastructure, containers, and Kubernetes.
- Experience with distributed storage such as NFS, HDFS, S3, or comparable cloud storage technologies.
- Hands-on troubleshooting across Linux, applications, networks, APIs, and distributed service dependencies.
- Working knowledge of Voice/UC or real-time communications concepts, plus experience using metrics, logs, traces, and service-level indicators to diagnose production issues.
Nice to have
- Experience supporting or integrating AI-enabled services, APIs, or workflows.
- Familiarity with speech and voice AI, machine-learning services, or LLM-based applications.
Culture & Benefits
- Primarily remote work with occasional office visits in Bristol or London.
- Teamwork, transparency, accountability, and cross-functional collaboration.
- Opportunities to work on cloud communications, real-time systems, and AI-enabled services.
- Full-time employment with an equal-opportunity and inclusive workplace.
Hiring process
- Application submission.
- Application review.
- Interview followed by the hiring decision.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
14 дней назад
Sr. Site Reliability Engineer / SWE
10 дней назад
Director of Platform Engineering (AI)
DeepL
14 дней назад
Site Reliability Engineer (AI)
12 дней назад
Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)
14 дней назад
Site Reliability Engineer (SRE)
10 дней назад