4 дня назад
Senior Site Reliability Engineer (AI Agents & Automation)
147 600 - 221 400$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI Agents & Automation) (Kubernetes/AWS/Azure): Building and operating reliable cloud infrastructure, observability systems, and autonomous AI agents for large-scale enterprise applications with an accent on Kubernetes, SLO-driven reliability, and production automation. Focus on designing agentic SRE workflows, resolving distributed-systems failures, and improving scalability, availability, and safe delivery through CI/CD.
Location: US Remote
Salary: $147,600–$221,400 USD annually for Zone 1 locations; $137,900–$206,900 USD annually for other US locations. Additional compensation may include an annual bonus, equity, and benefits.
Company
develops cloud-based business software that helps companies operate more efficiently.
What you will do
- Participate in on-call rotations, diagnose production issues, and perform root-cause analysis and remediation.
- Design and maintain observability dashboards, alerting, SLIs, SLOs, and error-budget practices.
- Operate and improve Kubernetes-based compute infrastructure across AWS and Azure.
- Build and operate autonomous AI agents that monitor infrastructure, remediate issues, and automate repetitive SRE work.
- Partner with product engineering teams on architecture, scalability, availability, performance, and reliability decisions.
- Maintain runbooks and documentation, contribute to CI/CD pipelines, and promote reliability best practices.
Requirements
- Strong hands-on Kubernetes experience.
- Practical experience building or operating autonomous AI agents for infrastructure and reliability work, including agent design, MCP, context management, authorization, and guardrails.
- Experience with SLIs, SLOs, error budgets, distributed systems, and production troubleshooting.
- Cloud engineering and networking experience with AWS or Azure, including subnetting and IP addressing.
- Deep experience with an observability stack such as OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch.
- Strong programming skills and 8–10+ years of relevant hands-on experience; .NET/ASP.NET, Python, or Java experience is suitable.
Nice to have
- Database experience.
- Experience with GitHub Actions, TeamCity, Azure DevOps, or GitLab CI.
Culture & Benefits
- Flexible time off, autonomous work support, learning and development opportunities, and leadership training.
- Recognition programs, peer-nominated awards, and career development support.
- Company-paid medical, dental, and vision coverage, with FSA, HSA, 401(k) matching, and telehealth options.
- Parental leave, fertility, surrogacy, and adoption support, plus pet insurance and financial planning resources.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
7 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
7 дней назад
Senior Site Reliability Engineer (Kubernetes)
170 000 - 185 000$
Vapi
11 дней назад
Senior Site Reliability Engineer (AI)
280 000 - 314 000$
Replit
5 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
5 дней назад