49 минут назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI): Maintaining and improving the production health of a cloud platform used for digital investigations with an accent on incident response, observability, and safe production changes. Focus on AI-assisted anomaly detection, cross-signal troubleshooting, zero-downtime upgrades, and coordinating complex incidents across support, R&D, and DevOps.
Location: US - New York City
Company
develops an AI-powered digital investigation platform that helps public safety organizations, intelligence agencies, and businesses lawfully access, analyze, and share digital evidence while preserving data privacy.
What you will do
- Own the full production incident lifecycle, including detection, severity assessment, investigation, resolution, and post-incident reviews.
- Coordinate major incident response across TCS support, R&D, and DevOps, developing toward an incident-commander role.
- Improve monitoring and observability through anomaly detection, log/metric/trace correlation, alert tuning, and AI-assisted triage.
- Plan and execute rolling upgrades, patches, and safe deployments with minimal downtime, including rollback validation.
- Maintain production runbooks, architecture documentation, and the known-issues knowledge base.
- Participate in an on-call rotation supporting production uptime SLOs.
Requirements
- 3–5 years of experience in production support, SRE, or operations.
- Hands-on experience operating and troubleshooting existing AWS infrastructure.
- Experience with monitoring and observability tools such as Datadog, CloudWatch, or Grafana, including investigation through logs, metrics, and traces.
- Working knowledge of Linux administration, basic networking, and containerized environments such as Kubernetes.
- Comfort with Python or Bash scripting and experience with rolling upgrades, patching, or zero-downtime production deployments.
- Strong written and verbal English is required for cross-team communication, handoffs, and documentation.
Nice to have
- Exposure to Terraform or other infrastructure-as-code tools.
- Basic CI/CD familiarity.
- Interest in applying AI/ML to monitoring, alerting, and incident triage.
Culture & Benefits
- Collaborative work across outsourced support, R&D, and DevOps teams.
- Blameless incident reviews focused on implementing corrective actions.
- Participation in an on-call rotation and production upgrade windows.
- Focus on protecting lives, accelerating justice, and preserving data privacy.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (AWS)
7 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
1 день назад
Site Reliability Engineer (AI/ML)
6 дней назад
Senior Site Reliability Engineer (AWS/AI)
140 000 - 160 000$
6 дней назад
Senior Site Reliability Engineer, Platform Infrastructure (AWS/AI)
5 дней назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$