4 часа назад
Senior Site Reliability Engineer (ML)
200 000 - 225 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (ML): Building and operating scalable infrastructure for data pipelines, machine learning workloads, and real-time analytics systems with an accent on reliability, observability, and operational excellence. Focus on automating CI/CD and deployment workflows, managing Kubernetes-based cloud infrastructure, and solving complex performance, monitoring, and incident-response challenges.
Location: On-site in Dunwoody, Georgia, Atlanta Perimeter
Salary: $200,000–$225,000 per year, plus equity
Company
builds a core platform for enterprise data-intensive workloads, including data pipelines, machine learning, and real-time analytics.
What you will do
- Design, build, and maintain scalable infrastructure for real-time analytics and machine learning workloads.
- Improve reliability and performance through automation, observability, capacity planning, and proactive system optimization.
- Own CI/CD pipelines, deployment automation, rollback mechanisms, and configuration management.
- Implement monitoring, alerting, SLOs, runbooks, and incident response processes, including on-call rotations.
- Collaborate with engineering and data science teams to improve platform performance and operational reliability.
- Drive security, compliance, post-incident analysis, and continuous infrastructure improvement.
Requirements
- 8+ years of experience in SRE, DevOps, or infrastructure engineering.
- 5+ years of experience in datacenter operations and/or systems and network administration.
- Experience with Docker, Kubernetes, Linux systems, networking, SSH, shell, and performance tuning.
- Experience with infrastructure as code and scripting using Terraform, Ansible, Bash, and/or Python.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, ELK, or OpenTelemetry.
- Strong communication skills and the ability to work cross-functionally in an on-site environment.
Nice to have
- Experience with AWS and other cloud-managed services.
- Experience supporting data platforms using Spark, Airflow, or Kafka.
- Familiarity with cloud-native security practices and SOC 2 or other high-compliance environments.
- Experience with GitHub Actions, ArgoCD, or similar CI/CD tools.
Culture & Benefits
- Ownership of mission-critical infrastructure supporting enterprise solutions.
- Opportunity to influence platform scaling, deployment, and incident management.
- Engineering culture focused on curiosity, accountability, performance, and impact.
- Market-based compensation with transparent, data-driven salary practices.
- Equity program for new hires.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior DevOps / Site Reliability Engineer (SRE) (Cybersecurity)
165 000 - 215 000$
2 дня назад
Site Reliability Engineer, Tech Lead (AI)
2 дня назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
1 день назад
Site Reliability Engineer II (Azure)
6 дней назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$
1 день назад