39 минут назад
Principal Engineer - Data Engineering (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Engineer - Data Engineering (AI): Building ML-ready scientific data pipelines for experimental measurement, sensor, and operational data with an accent on feature engineering, data quality, versioning, and lineage. Focus on reproducible training datasets, drift detection, streaming sensor ingestion, synthetic data infrastructure, and MLOps data layers integrated with AWS-based AI platforms.
Location: Singapore office
Company
builds storage systems and data infrastructure for AI-driven applications, hyperscale data centers, cloud platforms, and enterprise infrastructure.
What you will do
- Build and maintain versioned feature engineering pipelines that transform engineering, sensor, and operational data into ML-ready feature sets.
- Design data quality frameworks covering completeness, schema consistency, statistical stability, and label accuracy across AI training datasets.
- Implement data versioning, lineage tracking, reproducible training workflows, and deployment drift detection.
- Maintain data contracts, access controls, retention practices, and governance for AI data assets.
- Contribute to real-time sensor ingestion and synthetic data pipelines, progressively taking operational ownership over 6–12 months.
- Build the training dataset registry, feature store, and model input validation layer for the AI platform.
Requirements
- Bachelor’s or Master’s degree in AI, Computer Science, Data Engineering, Electrical Engineering, Applied Mathematics, or a related field.
- Fresh graduate to 1 year of experience, with demonstrated academic, personal, or internship project experience building end-to-end data pipelines.
- Strong Python and SQL skills, including complex queries, window functions, and data transformation logic.
- Knowledge of batch pipeline architecture, reliability, schema management, fault tolerance, and orchestration with Airflow, Prefect, or AWS Glue Jobs.
- Understanding of data quality, data versioning, lineage, ML data lifecycles, train/test leakage, and label quality impacts.
Nice to have
- Experience with feature stores such as Feast or Tecton, synthetic data generation, dbt, DLT, or annotation platforms.
- Knowledge of MES/LIMS integrations, active learning data loops, or data lakehouse technologies such as Iceberg, Dremio, AWS Glue, AWS Lake Formation, or Redshift.
Culture & Benefits
- Work in an inclusive environment focused on diversity, belonging, respect, and contribution.
- Build real-world data infrastructure supporting AI systems and global-scale manufacturing.
- Accessibility support is available throughout the application and hiring process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →