3 дня назад
Sr. Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. Site Reliability Engineer (AI/Kubernetes): Designing and operating highly available cloud-native production systems with an accent on Kubernetes, Terraform, observability, and automated infrastructure. Focus on AI-assisted incident investigations, reliable triage and root cause analysis, CI/CD delivery, autoscaling, and performance and cost optimization.
Location: Office - Boise, United States
Company
develops cloud-native analytics systems and applications.
What you will do
- Design, build, and maintain highly available, scalable, and reliable production systems.
- Define and manage SLIs, SLOs, and SLAs to improve system reliability.
- Provision and operate cloud infrastructure with Terraform and Amazon EKS.
- Maintain monitoring, logging, and alerting with Prometheus, Grafana, Dynatrace, and OpenSearch.
- Lead on-call incident response, production troubleshooting, root cause analysis, and reliability improvements.
- Build AI-assisted investigation workflows, CI/CD pipelines, self-healing mechanisms, and capacity and cost optimization practices.
Requirements
- 7+ years of experience in Site Reliability Engineering or Platform Engineering.
- Bachelor’s degree in computer science or a related field, or equivalent practical experience.
- Strong expertise in Terraform, Infrastructure as Code, AWS, and Amazon EKS.
- Hands-on experience with monitoring, logging, observability, incident management, on-call support, and root cause analysis.
- Proficiency in Python, Java, Go, or Bash, plus experience with Agile development and CI/CD pipelines.
- Experience with autoscaling, performance tuning, cost optimization, high-pressure production troubleshooting, and AI-assisted automation.
Nice to have
- Docker and Linux administration experience.
- Build systems and dependency management with Maven, Gradle, or npm.
- Experience with AWS services such as Cognito, WAF, Elasticsearch, SNS, SQS, S3, or Systems Manager.
- Database infrastructure knowledge, including RDS, MySQL, or SQL Server.
- Cloud or Kubernetes certifications.
Culture & Benefits
- Full-time office-based work in Boise.
- Participation in an on-call rotation supporting production systems.
- Collaboration with development teams to improve application reliability, scalability, and deployment processes.
- Focus on automation, operational excellence, security, compliance, and reducing manual toil.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer
90 000 - 110 000$
10 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
Replit
10 дней назад
Site Reliability Engineer
210 000 - 275 000$
4 дня назад
Site Reliability Engineering (SRE) Manager
106 000 - 130 600$
Replit
10 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
9 дней назад
Senior Site Reliability Engineer (AI)
185 500 - 232 000$