6 дней назад
Senior Data Engineer, Data Governance Lead (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Data Engineer, Data Governance Lead (AI): Building enterprise data quality, discovery, governance frameworks, and large-scale batch and real-time data pipelines with an accent on AI/ML-based anomaly detection, metadata-driven lineage, and cloud data platforms. Focus on designing petabyte-scale distributed systems, processing 10TB+ daily volumes and millions of events per minute, and leading cross-functional governance strategy.
Location: Austin, United States; hybrid work with office attendance generally required Monday through Thursday and up to one remote day per week
Company
operates a TV streaming platform that connects consumers, content publishers, and advertisers across the TV ecosystem.
What you will do
- Lead the architecture, design, and development of enterprise data quality, discovery, and governance frameworks.
- Establish governance standards, policies, tooling, and processes for accurate, compliant, and discoverable data assets.
- Build DataHub integrations and extend data quality frameworks such as Deequ and Great Expectations.
- Develop AI/ML solutions for data quality, anomaly detection, and automated data discovery.
- Design batch and real-time Spark pipelines processing more than 10TB daily and millions of events per minute.
- Lead engineers and consultants, plan roadmaps and sprints, and collaborate across product, analytics, engineering, and data science teams.
Requirements
- Master’s degree in Data Science, Computer Science, or a related field with 6 years of relevant experience, or a bachelor’s degree with 8 years of progressive experience.
- At least 1 year of experience leading architecture and governance initiatives, building enterprise data quality and discovery frameworks, and managing cross-functional delivery.
- At least 1 year of experience with AI/ML for data quality, anomaly detection, or data discovery, plus DataHub, Deequ, or Great Expectations.
- At least 5 years of experience with data modeling, SQL, Airflow, Python, Java or Scala, Spark, HDFS, YARN, Hive, Kafka, Flink, Elasticsearch, Grafana, and Presto.
- At least 5 years of experience developing distributed pipelines and data warehouses in AWS or GCP, including Spark on Kubernetes, CI/CD, serverless computing, and containerized deployments.
- Experience supporting large-scale data volumes, improving event logging, mentoring global engineering teams, and driving governance roadmaps.
Culture & Benefits
- Collaborative, fast-paced environment focused on practical innovation and customer delivery.
- Benefits may include healthcare, life and disability coverage, commuter support, retirement options including 401(k), and statutory or voluntary local benefits.
- Global mental health and financial wellness resources are available.
- Time off is provided according to local leave policies and personal needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Sr. Data & Software Engineer (AI)
90 000 - 132 000$
6 дней назад
Senior Data Engineer (AI)
11 дней назад
AI Data Engineer
152 000 - 186 000$
VIA
12 дней назад
Senior Data Analytics Engineer (AI)
140 000 - 160 000$
6 дней назад
Data Engineer (AI)
89 886 - 175 444$
6 дней назад