4 дня назад
Data Engineer
100 000 - 120 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Engineer (Spark/AWS): Building and operating production ETL pipelines on EMR with Airflow, using SQL and Python to deliver reliable datasets and dimensional data models with an accent on data quality, performance, and cost efficiency. Focus on designing recoverable workflows, optimizing Spark jobs, automating reusable platform patterns, and translating analyst and stakeholder needs into production data products.
Location: On-site in Irvine, California, United States
Base salary: $100,000–$120,000 per year
Company
develops networking devices and smart home products for customers in more than 170 countries.
What you will do
- Build, operate, monitor, and recover ETL pipelines on AWS EMR with Spark and Airflow.
- Write production SQL and Python, including PySpark, pandas, and boto3, for ingestion, transformation, and analytics datasets.
- Implement data quality checks and validate results before releasing pipelines.
- Develop fact and dimension tables according to warehouse layering and star-schema conventions.
- Optimize runtime and cloud costs, including Spark jobs and AWS resources.
- Partner with analysts and business stakeholders to turn requirements into production datasets.
Requirements
- 2–4 years of hands-on data development in a production environment.
- Strong SQL skills, including window functions, complex joins, incremental loads, and query performance analysis.
- Production Python experience for ETL and tooling, with PySpark, pandas, and boto3.
- Experience with Spark on AWS EMR, Databricks, Glue, or an equivalent cloud platform.
- Production scheduler experience, preferably Airflow, including DAGs, dependencies, retries, and safe reruns.
- Knowledge of dimensional modeling, AWS S3 with Parquet or ORC, Linux, Git, and effective AI-assisted problem solving. A bachelor's degree in computer science, information systems, or a related field, or equivalent practical experience, is required.
Nice to have
- Experience with DataX, Sqoop, Debezium, Fivetran, or similar ingestion and CDC tools.
- Experience with StarRocks, Doris, ClickHouse, Redshift, or another OLAP/MPP engine.
- Spark performance tuning, AWS cost optimization, lakehouse formats, streaming, dbt, or data quality frameworks.
- Experience with QuickSight or another BI tool and evidence of self-directed learning.
Culture & Benefits
- Fully paid medical, dental, and vision insurance, with partial dependent coverage.
- 401(k) contributions and health and wellness benefits, including a free gym membership.
- Bi-annual performance reviews and annual pay increases.
- Free snacks and drinks and quarterly team-building events.
- Visa sponsorship is not available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Mid-Level/Lead Data Engineer (Python, AWS, Spark)
85 500 - 130 000$
6 дней назад
Data Engineer (AI)
89 886 - 175 444$
4 дня назад
Data Engineer, AI Enablement
84 500 - 162 000$
10 дней назад
Data Engineer II (Azure)
100 411 - 128 604$
5 дней назад
Data Engineer (AI)
160 000 - 190 000$
6 дней назад
Data Engineer
125 000 - 170 000$