1 день назад
Data Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Engineer (AI) (Spark/Scala/AWS): Building and operating terabyte- and petabyte-scale data pipelines for AI-powered consumer intelligence with an accent on Spark performance, data quality, and production reliability. Focus on tuning distributed applications, debugging large-scale data systems, integrating Gen AI into engineering workflows, and supporting on-call operations.
Location: United States; remote-first with occasional travel for company and client events. is headquartered in Reston, Virginia, with offices in New York City and Washington, D.C.
Company
provides AI-powered consumer data, intelligence, and SaaS technology for marketers through its Ignite platform and Data-as-a-Service offerings.
What you will do
- Design, develop, and maintain ETL/ELT pipelines using Spark and Scala on AWS EMR across S3 and Snowflake.
- Tune Spark applications for performance and cost at multi-terabyte and petabyte scale.
- Take product requirements from design and planning through production delivery in partnership with senior engineers and product management.
- Monitor pipeline health with Grafana, investigate data-quality issues, and resolve production problems.
- Write clean, testable code with unit and integration tests, and use Gen AI tools across development and monitoring.
- Participate in code reviews, technical design discussions, sprint planning, on-call rotations, and incident response.
Requirements
- Approximately five or more years of experience in software engineering, data engineering, or a related field.
- At least three years of hands-on experience with Spark and Scala, including DataFrame and Dataset APIs.
- Proven experience tuning Spark applications at multi-terabyte or petabyte scale and debugging production big-data failures.
- Strong relational database experience and knowledge of data structures, data modeling, solution architecture, and the full software development lifecycle.
- Working knowledge of AWS EMR, S3, Lambda, Kafka, Snowflake, Grafana, Hadoop, Elastic Stack, and Docker.
- Genuine enthusiasm for applying generative AI to pipeline development, system monitoring, and developer productivity.
Nice to have
- Experience building and maintaining big-data pipelines with Gen AI.
- Experience with probabilistic data structures and high-cardinality data systems at scale.
- Bachelor’s degree in computer science, computer engineering, or equivalent experience.
Culture & Benefits
- Remote-first work with flexibility and intentional in-person collaboration when needed.
- Occasional travel for company and client events.
- 401(k) matching and an Open PTO policy.
- Comprehensive benefits for employees and their families.
- Opportunity to work on AI-powered marketing data and identity solutions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →