13 дней назад
Member of Technical Staff, ML Systems
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, ML Systems (CUDA/GNN): Building the internal ML platform that takes models from data and training through governed, reproducible releases across cloud and customer environments with an accent on GPU compute, distributed training, data materialization, and model lineage. Focus on designing CUDA kernels, improving multi-node training throughput, and ensuring every release can be audited and reproduced from its registered data, code, and configuration.
Location: Remote from São Paulo, Brazil
Company
builds machine learning systems and an internal platform for developing, governing, and deploying models across cloud and customer environments.
What you will do
- Build CUDA kernels and compute primitives for training and serving graph neural networks.
- Evolve the Monad sampler and distributed-training library, including neighbor sampling and training performance improvements.
- Define binary data formats and own materializations and feature backfills for training and evaluation.
- Establish data contracts and consumption requirements with teams producing customer and proprietary datasets.
- Build experiment tracking, checkpointing, evaluation, model registry, lineage, and versioning infrastructure.
- Define release gates and ensure production, batch, and on-premise models have governed, reproducible release records.
Requirements
- Strong systems engineering skills and production-quality Python.
- Experience with distributed training, such as Ray or PyTorch Distributed, and multi-node GPU workloads.
- Experience with columnar data formats and large-scale data materialization.
- Familiarity with experiment tracking, model registries, evaluation, and reproducibility tooling.
- Product mindset and ability to build internal platforms for researchers and platform engineers; data science experience is not required.
Nice to have
- CUDA kernel development or GPU performance optimization.
- Graph neural networks or graph sampling at scale.
- Lance, Arrow, or other columnar and indexed storage formats.
- Multi-cloud GPU compute, such as SkyPilot.
- Model governance or audit experience in financial services or other regulated environments.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →