5 дней назад
Big Data Development Engineer Intern (AI/Data Governance)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Big Data Development Engineer Intern (AI/Data Governance): Building AI-assisted capabilities for data model governance, metric semantic automation, and data quality management with an accent on SQL, Python, data lineage, and LLM applications. Focus on designing and deploying an end-to-end MVP, detecting data and metric issues, evaluating model effectiveness, and generating automated remediation recommendations.
Location: Hong Kong SAR; 5 days per week on-site
Company
is a cryptocurrency exchange and digital financial platform providing trading, payments, wealth management, custody, institutional services, and Web3 products.
What you will do
- Research AI and data governance approaches, including metadata and lineage platforms, semantic layers, data quality frameworks, and natural-language metric querying.
- Build AI-assisted data model review capabilities for naming, warehouse layering, duplicate detection, asset value analysis, and refactoring recommendations.
- Extract and standardize metric definitions from SQL, lineage, and documentation while detecting conflicting or redundant metric implementations.
- Automate data quality rules, anomaly detection, lineage-based root cause analysis, alert grading, and remediation recommendations.
- Deliver at least one end-to-end MVP on a real data domain, measure its results, and iterate based on findings.
- Document solution designs and evaluation methods, present results to data stakeholders, and create reusable frameworks.
Requirements
- Undergraduate or graduate student in Computer Science, Data Science, Statistics, or a related field.
- Solid SQL and Python skills, with knowledge of dimensional modeling, layered data warehouse architecture, metadata, and data lineage.
- Hands-on experience developing LLM applications using prompt engineering, RAG, or agent and tool-calling frameworks, with the ability to evaluate system effectiveness.
- Structured problem-solving skills, including defining measurable success criteria for ambiguous data governance problems.
- Minimum 3-month internship commitment and availability for 5 days per week on-site in Hong Kong SAR.
- Fluent Mandarin is required. Ability to read technical materials in English and clear written communication skills are required.
Nice to have
- Experience with Spark, Flink, Hive, or StarRocks.
- Experience with DataHub, OpenMetadata, Atlas, dbt, Great Expectations, Deequ, or metric and semantic-layer tools.
- Fluent English.
Culture & Benefits
- Study Growth Fund supporting professional development and continuous learning.
- Regular team-building activities, workshops, and internal events.
- Collaboration with an international team.
- Career advancement and internal mobility opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →