Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Research Data Engineer (AI) (Python/ML): Building data pipelines, datasets, and tooling that turn multimodal agent research into reliable training and evaluation workflows with an accent on large-scale data processing, quality assurance, and distributed storage. Focus on designing annotation and synthetic data systems, monitoring dataset reliability, and supporting LLM/VLM training with reproducible data foundations.
Location: Vienna, Austria
Company
Canva is a design platform developing AI products that help millions of people create with confidence.
What you will do
- Design and build data pipelines for collecting, filtering, deduplicating, formatting, and versioning text, image, and multimodal datasets.
- Build infrastructure for large-scale data loading, storage, and retrieval using AWS, S3, distributed systems, and streaming pipelines.
- Translate research requirements into data specifications and collaborate with research scientists on experiments and benchmarks.
- Develop dataset construction tooling, including human annotation workflows, synthetic data generation, and preference data collection for RLHF/DPO-style training.
- Build validation, monitoring, testing, and documentation systems to ensure dataset quality, provenance, reproducibility, and pipeline reliability.
- Identify data bottlenecks, improve code quality, and contribute solutions to research roadmaps.
Requirements
- Strong Python software engineering skills and experience building production-grade data pipelines and ML DevOps systems.
- Practical experience with prompt engineering for reliable LLM/VLM outputs.
- Experience with large-scale distributed ML data workflows, data versioning, tokenization, batching, and sharding.
- Experience with annotation tooling, human-in-the-loop data collection, and large datasets in AWS and distributed storage systems.
- Understanding of data requirements for LLM/VLM fine-tuning, along with strong communication and collaboration skills.
- Ability to take ownership of ambiguous data problems and iterate quickly with research teams.
Nice to have
- Experience with preference data collection for RLHF or reward modeling.
- Familiarity with multimodal data such as image-text pairs, video, and design assets.
- Experience building synthetic data generation pipelines with LLMs.
- Background in data quality metrics, monitoring systems, dataset releases, or ML benchmarks.
Culture & Benefits
- Full-time permanent employment.
- Choice in how and where to work within the available work arrangement.
- Collaboration with research, product, and platform teams.
- Virtual interviews.
- Reasonable adjustments are available during the interview process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →