14 часов назад
AWS Lakehouse Data Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AWS Lakehouse Data Engineer (AWS/Python/PySpark): Building and operating AWS-native lakehouse platforms and production data pipelines with an accent on Apache Iceberg, data governance, reliability, and cost optimization. Focus on implementing ACID transactions, schema evolution, lineage, quality controls, Infrastructure as Code, and CI/CD across scalable analytics environments.
Location: Remote – offsite; role listed in McLean, Virginia, United States
Company
delivers outsourced services and workforce solutions across North America.
What you will do
- Build and operate batch and streaming data pipelines from APIs, databases, files, event streams, and external partners.
- Design and optimize ETL/ELT workflows with Python and PySpark, including incremental processing, CDC, data contracts, and schema validation.
- Design and manage an AWS-native lakehouse on Amazon S3 using Apache Iceberg, Parquet, Athena, EMR, Glue, and Redshift.
- Implement ACID transactions, schema evolution, snapshot isolation, time travel, data governance, lineage, access controls, and metadata management.
- Improve reliability, quality, performance, and cost through testing, orchestration, monitoring, partitioning, caching, lifecycle policies, and SLA/SLO metrics.
- Automate AWS provisioning with IaC, develop CI/CD pipelines, standardize environment promotion, and maintain engineering documentation.
Requirements
- Bachelor’s degree in a relevant technical field or four years of equivalent practical experience.
- Six years of relevant professional experience.
- Hands-on experience building AWS-native data lake or lakehouse architectures on Amazon S3.
- Strong production experience with Python, PySpark, advanced SQL, data modeling, transformation, and performance tuning.
- Experience with Apache Iceberg, AWS data governance, cataloging, lineage, access control, IAM, KMS, secrets management, logging, and network security.
- Experience with IaC, CI/CD for data workflows, distributed workload troubleshooting, cost management, and cross-functional collaboration.
Nice to have
- Experience with Databricks, Delta Lake, and migration to AWS-native services.
- Familiarity with Step Functions, MWAA, Kinesis, DMS, Lambda, MSK, Git, Terraform, CloudFormation, Jenkins, Docker, or GitHub Actions.
- Knowledge of QuickSight, Tableau, Power BI, AI-assisted coding tools, graph modeling, ontology, taxonomy, entity resolution, or hybrid retrieval.
Culture & Benefits
- Remote work model with cross-functional collaboration.
- Eligible employees may receive medical, dental, vision, spending account, life insurance, and voluntary plan options.
- Eligible employees may participate in a 401(k) plan.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Sr. Data Engineer (AI)
143 065 - 241 598$
5 дней назад
Senior Data Engineer (AI)
152 000 - 190 000$
3 дня назад
Senior Data Engineer (AI)
119 514 - 164 899$
4 дня назад
Senior Data Engineer, Finance (AWS)
161 500 - 283 900$
5 дней назад
Manager, Analytics Engineering
165 000 - 195 000$
6 дней назад