обновлено 1 день назад
Senior Data Reliability Engineer (AWS)
105 700 - 149 275$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Data Reliability Engineer (AWS): Operating and improving production data pipelines and AWS-based data platform services with an accent on reliability, observability, incident response, and data SLAs. Focus on troubleshooting distributed Spark/EMR and warehouse systems, performing root cause analysis, automating operational workflows, and building resilient disaster recovery processes.
Location: Remote - Nationwide, United States. Applicants must be authorized to work for any employer in the U.S.; visa sponsorship, including CPT/OPT sponsorship, is not available.
Base salary: $105,700–$149,275 per year, with eligibility for a bonus program.
Company
provides financial services focused on helping customers achieve financial freedom.
What you will do
- Own the reliability and stability of production data pipelines and AWS data platform services.
- Diagnose pipeline failures, delays, data quality issues, and performance problems across distributed systems such as Spark, EMR, Redshift, and streaming platforms.
- Lead or support incident response, including triage, mitigation, root cause analysis, and durable remediation.
- Define and improve data SLAs covering freshness, latency, and completeness.
- Design monitoring, alerting, observability, automation, disaster recovery, backup validation, and recovery workflows.
- Partner with engineering teams and maintain runbooks, standard operating procedures, and operational documentation.
Requirements
- At least 5 years of experience working with production data platforms in AWS environments.
- Experience building and operating production data pipelines, including troubleshooting real-world failures.
- Strong Python and SQL skills in real data systems.
- Hands-on experience troubleshooting distributed data processing systems such as Spark/EMR, Redshift, and streaming systems.
- Experience with AWS data services including EMR, Redshift, DynamoDB, S3, or similar technologies.
- Experience handling production incidents, performing root cause analysis, and resolving ambiguous operational issues.
Nice to have
- Experience with backfills, reprocessing, late-arriving data, and incomplete data.
- Experience improving observability and alerting for data systems.
- Exposure to Kafka, Kinesis, CDC patterns, or other event-driven systems.
- Experience with disaster recovery, backup validation, and resiliency testing.
- Strong communication with technical and non-technical stakeholders during incidents.
Culture & Benefits
- Flexible remote work environment with reliable high-speed wired internet required.
- Medical, dental, vision, and life insurance.
- 401(k) plan with company matching contributions of up to 6%.
- Tuition reimbursement of up to $5,250 per year.
- Paid time off, company holidays, floating holidays, volunteer time, parental leave, disability coverage, and FMLA programs.
- Inclusive business resource groups and a business-casual work environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (GovCloud)
117 000 - 209 330$
5 дней назад
Sr. Site Reliability Engineer (Healthcare)
125 000 - 145 000$
6 дней назад
Senior Data Platform Engineer (Web3)
126 000 - 180 000$
7 дней назад
Senior Site Reliability Engineer (Kubernetes)
125 000 - 145 000$
6 дней назад
Senior DevOps Engineer (AWS/Kubernetes)
140 000 - 165 000$
2 дня назад
AWS Cloud Engineer
100 000 - 150 000$