Назад
Company hidden
14 часов назад

AWS Lakehouse Data Engineer

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AWS Lakehouse Data Engineer (AWS/Python/PySpark): Building and operating AWS-native lakehouse platforms and production data pipelines with an accent on Apache Iceberg, data governance, reliability, and cost optimization. Focus on implementing ACID transactions, schema evolution, lineage, quality controls, Infrastructure as Code, and CI/CD across scalable analytics environments.

Location: Remote – offsite; role listed in McLean, Virginia, United States

Company

hirify.global delivers outsourced services and workforce solutions across North America.

What you will do

  • Build and operate batch and streaming data pipelines from APIs, databases, files, event streams, and external partners.
  • Design and optimize ETL/ELT workflows with Python and PySpark, including incremental processing, CDC, data contracts, and schema validation.
  • Design and manage an AWS-native lakehouse on Amazon S3 using Apache Iceberg, Parquet, Athena, EMR, Glue, and Redshift.
  • Implement ACID transactions, schema evolution, snapshot isolation, time travel, data governance, lineage, access controls, and metadata management.
  • Improve reliability, quality, performance, and cost through testing, orchestration, monitoring, partitioning, caching, lifecycle policies, and SLA/SLO metrics.
  • Automate AWS provisioning with IaC, develop CI/CD pipelines, standardize environment promotion, and maintain engineering documentation.

Requirements

  • Bachelor’s degree in a relevant technical field or four years of equivalent practical experience.
  • Six years of relevant professional experience.
  • Hands-on experience building AWS-native data lake or lakehouse architectures on Amazon S3.
  • Strong production experience with Python, PySpark, advanced SQL, data modeling, transformation, and performance tuning.
  • Experience with Apache Iceberg, AWS data governance, cataloging, lineage, access control, IAM, KMS, secrets management, logging, and network security.
  • Experience with IaC, CI/CD for data workflows, distributed workload troubleshooting, cost management, and cross-functional collaboration.

Nice to have

  • Experience with Databricks, Delta Lake, and migration to AWS-native services.
  • Familiarity with Step Functions, MWAA, Kinesis, DMS, Lambda, MSK, Git, Terraform, CloudFormation, Jenkins, Docker, or GitHub Actions.
  • Knowledge of QuickSight, Tableau, Power BI, AI-assisted coding tools, graph modeling, ontology, taxonomy, entity resolution, or hybrid retrieval.

Culture & Benefits

  • Remote work model with cross-functional collaboration.
  • Eligible employees may receive medical, dental, vision, spending account, life insurance, and voluntary plan options.
  • Eligible employees may participate in a 401(k) plan.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →