4 дня назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI) (Cloud Communications): Building and operating reliable production infrastructure for distributed Voice/Unified Communications services and AI-enabled communication capabilities with an accent on observability, scalability, and real-time service quality. Focus on designing resilient Voice/UC and AI integrations, automating production-readiness checks, and solving complex failure, latency, jitter, and dependency-isolation challenges.
Location: Portugal; fully remote
Company
provides cloud communications and collaboration technology for businesses, including Voice and Unified Communications services with AI-powered capabilities.
What you will do
- Run and improve production environments by monitoring availability, system health, performance, and capacity.
- Build software, automation, and infrastructure systems for large distributed applications and services.
- Strengthen reliability through observability, incident response, post-incident improvements, SLIs, SLOs, and error budgets.
- Improve Voice/UC platforms and integrations, including real-time signaling, media flows, call quality, latency, jitter, packet loss, failover, and service availability.
- Operationalize AI-enabled communications capabilities such as speech recognition, transcription, summarization, intelligent routing, conversational assistance, and text-to-speech.
- Design resilient integrations with graceful degradation, dependency isolation, retry and fallback patterns, recovery procedures, and automated production-readiness validation.
Requirements
- Bachelor’s degree in computer science or a technical/scientific discipline, or equivalent practical experience.
- 4–7 years of experience in production operations, systems engineering, SRE/DevOps, CI/CD, software deployment, and production system maintenance.
- Experience with Agile, DevOps, CI/CD pipelines, infrastructure automation, monitoring, and observability.
- Experience with distributed systems, cloud infrastructure, containers, Kubernetes, and distributed storage such as NFS, HDFS, or S3.
- Hands-on troubleshooting across Linux, applications, networks, APIs, and distributed service dependencies.
- Working knowledge of Voice/UC or real-time communications, plus experience with AI-enabled services, APIs, or workflows and metrics, logs, traces, and service-level indicators.
Culture & Benefits
- Fully remote work arrangement for the Portugal-based role.
- Teamwork, transparency, accountability, and cross-functional collaboration.
- Opportunities to work with cloud communications, real-time services, and AI-enabled products.
- Equal opportunity employment and commitment to reasonable workplace accommodations.
Hiring process
- Application review followed by an interview.
- Successful candidates proceed to hiring.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →