5 дней назад
Senior Engineer Site Reliability (AWS/Data Operations)
105 700 - 149 275$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Engineer Site Reliability (AWS/Data Operations): Improving the reliability, observability, and operational maturity of cloud-based data platforms and production data pipelines with an accent on AWS, incident response, automation, and infrastructure as code. Focus on designing operational monitoring, diagnosing complex production failures, reducing alert noise, and building resilient recovery and remediation practices.
Location: Remote - Nationwide within the United States. Applicants must be authorized to work for any employer in the U.S.; visa sponsorship, including CPT/OPT sponsorship, is not available.
Base salary: $105,700–$149,275 per year, with potential participation in a bonus program.
Company
provides financial services focused on helping customers achieve financial freedom.
What you will do
- Own and improve the reliability, availability, performance, and operational health of production data platforms and pipelines.
- Monitor pipeline execution, dependencies, failures, recovery, delays, and downstream impact.
- Troubleshoot complex production issues across AWS, data platforms, pipelines, and supporting services.
- Design observability with Datadog and Splunk, including dashboards, alerting, logging, and service-health indicators.
- Lead incident response, service restoration, root-cause analysis, corrective actions, and operational runbook improvements.
- Automate monitoring, health checks, remediation, and repetitive support tasks with Python while improving resilience, capacity, cost efficiency, and operational maturity.
Requirements
- 5+ years of hands-on AWS experience supporting production environments.
- Experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Cloud Reliability Engineering.
- Strong practical knowledge of SRE principles, production operations, incident management, and root-cause analysis.
- Production experience with Amazon Redshift or similar enterprise data platforms and data pipelines.
- Hands-on experience with Datadog and Splunk, plus strong Python programming and working SQL knowledge.
- Strong understanding of Terraform and Infrastructure as Code, including reading, reviewing, and troubleshooting existing IaC.
Nice to have
- Production experience with Snowflake or multiple cloud-based data platforms.
- Experience designing observability for large-scale, distributed, or data-intensive systems.
- Experience with AIOps, intelligent automation, agentic operations, or AI-assisted incident management.
- Familiarity with Amazon CloudWatch, data-pipeline orchestration, failure recovery, and cloud-cost visibility.
- Experience establishing or maturing SRE practices within an engineering organization.
Culture & Benefits
- Flexible remote work environment with a focus on purpose, well-being, inclusion, and work-life balance.
- Medical, dental, vision, and life insurance.
- 401(k) plan with company matching contributions of up to 6% and additional retirement benefits.
- Paid time off, ten paid company holidays, floating holidays, paid parental leave, disability leave, and FMLA programs.
- Tuition reimbursement of up to $5,250 per year and 16 hours of paid volunteer time annually.
- PagerDuty on-call rotation with occasional weekend coverage.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Senior Site Reliability Engineer (AWS)
7 дней назад
Sr. Manager, Site Reliability
59 550 - 110 594GBP
Okta
10 дней назад
Staff Site Reliability Engineer, Federal (TS/SCI)
174 000 - 238 000$
9 дней назад
Staff Site Reliability Engineer (Remote, AWS)
170 000 - 210 000$
10 дней назад
Senior Site Reliability Engineer (AWS/AI)
140 000 - 160 000$
10 дней назад
Senior Site Reliability Engineer (AWS)
130 000 - 155 000$