Назад
Company hidden
обновлено 1 день назад

Senior Data Reliability Engineer (AWS)

105 700 - 149 275$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Data Reliability Engineer (AWS): Operating and improving production data pipelines and AWS-based data platform services with an accent on reliability, observability, incident response, and data SLAs. Focus on troubleshooting distributed Spark/EMR and warehouse systems, performing root cause analysis, automating operational workflows, and building resilient disaster recovery processes.

Location: Remote - Nationwide, United States. Applicants must be authorized to work for any employer in the U.S.; visa sponsorship, including CPT/OPT sponsorship, is not available.

Base salary: $105,700–$149,275 per year, with eligibility for a bonus program.

Company

hirify.global provides financial services focused on helping customers achieve financial freedom.

What you will do

  • Own the reliability and stability of production data pipelines and AWS data platform services.
  • Diagnose pipeline failures, delays, data quality issues, and performance problems across distributed systems such as Spark, EMR, Redshift, and streaming platforms.
  • Lead or support incident response, including triage, mitigation, root cause analysis, and durable remediation.
  • Define and improve data SLAs covering freshness, latency, and completeness.
  • Design monitoring, alerting, observability, automation, disaster recovery, backup validation, and recovery workflows.
  • Partner with engineering teams and maintain runbooks, standard operating procedures, and operational documentation.

Requirements

  • At least 5 years of experience working with production data platforms in AWS environments.
  • Experience building and operating production data pipelines, including troubleshooting real-world failures.
  • Strong Python and SQL skills in real data systems.
  • Hands-on experience troubleshooting distributed data processing systems such as Spark/EMR, Redshift, and streaming systems.
  • Experience with AWS data services including EMR, Redshift, DynamoDB, S3, or similar technologies.
  • Experience handling production incidents, performing root cause analysis, and resolving ambiguous operational issues.

Nice to have

  • Experience with backfills, reprocessing, late-arriving data, and incomplete data.
  • Experience improving observability and alerting for data systems.
  • Exposure to Kafka, Kinesis, CDC patterns, or other event-driven systems.
  • Experience with disaster recovery, backup validation, and resiliency testing.
  • Strong communication with technical and non-technical stakeholders during incidents.

Culture & Benefits

  • Flexible remote work environment with reliable high-speed wired internet required.
  • Medical, dental, vision, and life insurance.
  • 401(k) plan with company matching contributions of up to 6%.
  • Tuition reimbursement of up to $5,250 per year.
  • Paid time off, company holidays, floating holidays, volunteer time, parental leave, disability coverage, and FMLA programs.
  • Inclusive business resource groups and a business-casual work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →