Назад
Company hidden
49 минут назад

Site Reliability Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI): Maintaining and improving the production health of a cloud platform used for digital investigations with an accent on incident response, observability, and safe production changes. Focus on AI-assisted anomaly detection, cross-signal troubleshooting, zero-downtime upgrades, and coordinating complex incidents across support, R&D, and DevOps.

Location: US - New York City

Company

hirify.global develops an AI-powered digital investigation platform that helps public safety organizations, intelligence agencies, and businesses lawfully access, analyze, and share digital evidence while preserving data privacy.

What you will do

  • Own the full production incident lifecycle, including detection, severity assessment, investigation, resolution, and post-incident reviews.
  • Coordinate major incident response across TCS support, R&D, and DevOps, developing toward an incident-commander role.
  • Improve monitoring and observability through anomaly detection, log/metric/trace correlation, alert tuning, and AI-assisted triage.
  • Plan and execute rolling upgrades, patches, and safe deployments with minimal downtime, including rollback validation.
  • Maintain production runbooks, architecture documentation, and the known-issues knowledge base.
  • Participate in an on-call rotation supporting production uptime SLOs.

Requirements

  • 3–5 years of experience in production support, SRE, or operations.
  • Hands-on experience operating and troubleshooting existing AWS infrastructure.
  • Experience with monitoring and observability tools such as Datadog, CloudWatch, or Grafana, including investigation through logs, metrics, and traces.
  • Working knowledge of Linux administration, basic networking, and containerized environments such as Kubernetes.
  • Comfort with Python or Bash scripting and experience with rolling upgrades, patching, or zero-downtime production deployments.
  • Strong written and verbal English is required for cross-team communication, handoffs, and documentation.

Nice to have

  • Exposure to Terraform or other infrastructure-as-code tools.
  • Basic CI/CD familiarity.
  • Interest in applying AI/ML to monitoring, alerting, and incident triage.

Culture & Benefits

  • Collaborative work across outsourced support, R&D, and DevOps teams.
  • Blameless incident reviews focused on implementing corrective actions.
  • Participation in an on-call rotation and production upgrade windows.
  • Focus on protecting lives, accelerating justice, and preserving data privacy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →