Назад
Company hidden
обновлено 8 дней назад

Data Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
Germany
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Engineer (AI) (Python/GCP): Building the software backbone for an AI data platform that ingests, curates, catalogs, and serves terabytes of multi-sensor field data with an accent on ETL pipelines, metadata systems, and automated labeling workflows. Focus on designing searchable data tooling, integrating field and annotation data, and ensuring reliable delivery to ML workflows.

Location: Munich or Berlin, Germany

Company

hirify.global is a defence technology startup developing and manufacturing software-defined, scalable unmanned systems for NATO Allies and their partners.

What you will do

  • Design, build, and maintain Python software and ETL pipelines for flight data, rosbags, video, and external data deliveries.
  • Develop tools for video sub-sampling, frame extraction, dataset packaging, and self-service data access.
  • Build metadata databases, data catalogs, schemas, and lineage tracking for datasets, recordings, sensors, and labels.
  • Own technical workflows for data curation and labeling, coordinate with external annotation vendors, and implement automated quality assurance.
  • Develop backend APIs, dashboards, and data-access tools for AI and engineering teams.
  • Improve testing, typing, documentation, logging, CI/CD, access management, and automation across the data platform.

Requirements

  • Strong experience writing clean, typed, tested, and maintainable Python.
  • Proven experience designing, building, and operating ETL or data-ingestion pipelines.
  • Strong SQL and database design skills, preferably with PostgreSQL.
  • Hands-on experience with object storage such as GCP/GCS or S3, plus Docker and CI/CD fundamentals.
  • Comfort working in Unix/Linux environments with shell scripting and standard Unix tools.
  • Ability to collaborate cross-functionally and coordinate data labeling and curation with external vendors and non-technical stakeholders.

Nice to have

  • Experience with synthetic data, GenAI-assisted workflows, auto-labeling, or data augmentation.
  • Familiarity with DVC, LakeFS, FiftyOne, COCO, or large-scale data curation workflows.
  • Experience with BigQuery, Cloud Run, IAM, robotics data formats, or multimodal data such as video and lidar.

Culture & Benefits

  • Work in a young startup environment with evolving priorities and significant ownership.
  • Build data systems used directly by perception, autonomy, AI, and engineering teams.
  • Critical thinking, grit, adaptability, and learning ability are valued over matching every qualification.
  • Permanent, full-time employment.

Hiring process

  • The process emphasizes critical thinking, grit, and the ability to learn rather than a perfect checklist match.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →