3 дня назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI): Building and improving platform systems, CI/CD pipelines, developer tooling, and agent harnesses for large-scale, real-time conversational AI traffic with an accent on platform reliability, observability, and autonomous-agent infrastructure. Focus on designing scalable Kubernetes and cloud systems, improving deployment ergonomics, and managing incidents while shaping AI-native platform engineering patterns.
Location: Remote
Company
builds AI-powered conversation intelligence and AI agents for enterprise contact centers, using large language models and voice AI to handle customer interactions at scale.
What you will do
- Shape the design and implementation of platform engineering, site reliability, cloud infrastructure, developer experience, and cost-visibility domains.
- Build systems that reduce operational toil and keep production infrastructure available and operable under large-scale, real-time conversational AI traffic.
- Extend the agent harness with CI, sandboxes, guardrails, validation, and agent-first evaluation loops.
- Own and improve CI/CD pipelines and developer tooling, including build and test performance, deployment ergonomics, caching, and self-service paths for new services.
- Participate in the on-call rotation and incident management process; non-business-hours pages are rare.
Requirements
- 6+ years of experience in software development enablement roles.
- End-to-end ownership of CI/CD platforms, including architecture, caching, and developer self-service.
- Effective use of AI tools such as Claude and Cursor for coding, troubleshooting, and reasoning, with sound judgment about where not to use them.
- Experience making code changes in Node/TypeScript and using Python and Terraform for automation in a Kubernetes/Helm environment.
- Practical experience with observability, including logs, metrics, tracing, monitoring, alerting, incident management, and related tooling.
- Experience working effectively in fully remote teams.
Nice to have
- Experience building harnesses or platforms for autonomous AI agents.
- Production-at-scale experience with GCP.
- Experience with telephony and SIP architectures, particularly FreeSWITCH.
Culture & Benefits
- Distributed remote work with flexibility for personal life events and calendar management.
- Company-wide offsites and smaller team gatherings.
- Funding for conferences, books, courses, and other learning activities.
- Flexible vacation, paid sabbatical after five years, comprehensive benefits, and physical and mental wellness support.
- Competitive compensation and equity in a fast-growing AI company.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Engineering Manager, SRE (AI)
9 дней назад
Staff Site Reliability Engineer (GCP/Kubernetes)
5 дней назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
9 дней назад
Senior Site Reliability Engineer (Cloud-Native Infrastructure)
142 800 - 178 500$
4 дня назад
Site Reliability Engineer (AI)
106 029 - 118 503$
9 дней назад