Назад
Company hidden
4 дня назад

Principal Engineer - Data Engineering (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Engineer - Data Engineering (AI): Building versioned ML-ready scientific data pipelines for experimental measurement, sensor, and operational data with an accent on data quality, lineage, reproducibility, and feature engineering. Focus on designing streaming ingestion, synthetic data infrastructure, data contracts, and MLOps data layers for AI training and deployment.

Location: Singapore office

Company

hirify.global builds storage systems and data infrastructure for AI-driven applications, hyperscale data centers, cloud platforms, and enterprise systems.

What you will do

  • Build and maintain versioned feature engineering pipelines that transform engineering, sensor, and operational data into ML-ready feature sets.
  • Design data quality frameworks covering completeness, schema consistency, statistical distribution stability, and label accuracy.
  • Implement data versioning, lineage tracking, and drift detection to support reproducible model training and reliable deployment data.
  • Maintain data contracts, access controls, retention practices, and governance for AI data assets.
  • Contribute to real-time sensor data ingestion and progressively take operational ownership over 6–12 months.
  • Build infrastructure for synthetic data generation, training dataset registries, feature stores, and model input validation integrated with the AI platform.

Requirements

  • Bachelor’s or master’s degree in AI, Computer Science, Data Engineering, Electrical Engineering, Applied Mathematics, or a related field.
  • Fresh graduate to 1 year of experience, with demonstrated academic, personal, or internship projects involving end-to-end data pipelines.
  • Strong Python and SQL skills, including complex queries, window functions, and data transformation logic.
  • Understanding of batch pipeline architecture, reliability, schema management, fault tolerance, and orchestration with Airflow, Prefect, or AWS Glue Jobs.
  • Knowledge of data quality, data versioning, lineage, reproducibility, train/test leakage, and label quality risks in ML pipelines.

Nice to have

  • Experience with Feast or Tecton, synthetic data generation, dbt, or DLT.
  • Experience integrating annotation platforms such as Label Studio or CVAT.
  • Experience with MES/LIMS systems, active learning data loops, data lakehouse technologies, Iceberg, Dremio, AWS Lake Formation, or Redshift.

Culture & Benefits

  • Full-time position based in the Singapore office.
  • Inclusive environment focused on diversity, belonging, respect, and contribution.
  • Candidate accessibility accommodations are available throughout the hiring process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →