Назад
Company hidden
4 часа назад

Senior Site Reliability Engineer (ML)

200 000 - 225 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (ML): Building and operating scalable infrastructure for data pipelines, machine learning workloads, and real-time analytics systems with an accent on reliability, observability, and operational excellence. Focus on automating CI/CD and deployment workflows, managing Kubernetes-based cloud infrastructure, and solving complex performance, monitoring, and incident-response challenges.

Location: On-site in Dunwoody, Georgia, Atlanta Perimeter

Salary: $200,000–$225,000 per year, plus equity

Company

hirify.global builds a core platform for enterprise data-intensive workloads, including data pipelines, machine learning, and real-time analytics.

What you will do

  • Design, build, and maintain scalable infrastructure for real-time analytics and machine learning workloads.
  • Improve reliability and performance through automation, observability, capacity planning, and proactive system optimization.
  • Own CI/CD pipelines, deployment automation, rollback mechanisms, and configuration management.
  • Implement monitoring, alerting, SLOs, runbooks, and incident response processes, including on-call rotations.
  • Collaborate with engineering and data science teams to improve platform performance and operational reliability.
  • Drive security, compliance, post-incident analysis, and continuous infrastructure improvement.

Requirements

  • 8+ years of experience in SRE, DevOps, or infrastructure engineering.
  • 5+ years of experience in datacenter operations and/or systems and network administration.
  • Experience with Docker, Kubernetes, Linux systems, networking, SSH, shell, and performance tuning.
  • Experience with infrastructure as code and scripting using Terraform, Ansible, Bash, and/or Python.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, ELK, or OpenTelemetry.
  • Strong communication skills and the ability to work cross-functionally in an on-site environment.

Nice to have

  • Experience with AWS and other cloud-managed services.
  • Experience supporting data platforms using Spark, Airflow, or Kafka.
  • Familiarity with cloud-native security practices and SOC 2 or other high-compliance environments.
  • Experience with GitHub Actions, ArgoCD, or similar CI/CD tools.

Culture & Benefits

  • Ownership of mission-critical infrastructure supporting enterprise solutions.
  • Opportunity to influence platform scaling, deployment, and incident management.
  • Engineering culture focused on curiosity, accountability, performance, and impact.
  • Market-based compensation with transparent, data-driven salary practices.
  • Equity program for new hires.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →