обновлено 1 месяц назад
Machine Learning Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Machine Learning Engineer (Python/ML Infrastructure): Architecting and developing a scalable research platform for large-scale experimentation, model training, and simulation across on-premises HPC and multi-cloud environments with an accent on distributed GPU workloads, reproducibility, and observability. Focus on designing high-throughput training pipelines, improving resource scheduling and workload isolation, and building tools for feature engineering, dataset generation, and large-scale backtesting.
Location: Hong Kong, Hong Kong; hybrid working opportunities are available.
Company
Quantitative trading firm developing electronic trading infrastructure, machine learning systems, and high-performance technology for independent trading teams.
What you will do
- Architect and develop a scalable, reliable, observable, and reproducible machine learning research platform.
- Build infrastructure for large-scale experimentation, model training, and simulation across on-premises HPC and multi-cloud environments.
- Design and optimize distributed training pipelines for high-throughput, GPU-accelerated workloads.
- Improve experiment management, model versioning, artifact tracking, and data lineage.
- Develop tools for feature engineering, dataset generation, and large-scale backtesting.
- Lead improvements to compute efficiency, resource scheduling, workload isolation, and ML platform observability.
Requirements
- 2+ years of experience designing and building large-scale distributed systems, ideally for research or data-intensive workloads.
- Strong Python programming skills with a focus on clean, maintainable, and high-performance code.
- Experience operating applications on Linux-based HPC clusters and/or cloud platforms.
- Understanding of distributed computing, parallel processing, and resource management.
- Experience with GPU-based workloads and modern ML frameworks such as PyTorch, TensorFlow, or JAX.
- Experience optimizing data pipelines and handling large-scale structured and unstructured datasets, with strong troubleshooting and communication skills.
Nice to have
- Experience building internal ML platforms or research tooling at scale.
- Familiarity with experiment tracking, workflow orchestration, and model lifecycle management.
- Experience with Docker and Kubernetes.
- Exposure to quantitative finance, simulation systems, or latency- and performance-sensitive domains.
Culture & Benefits
- Collaborative, welcoming, and diverse workplace with minimal hierarchy and an emphasis on respectful teamwork.
- Generous paid time off policies and regional savings plans.
- Free breakfast, lunch, and snacks daily.
- In-office wellness experiences, wellness reimbursements, company-sponsored sports teams, and fitness events.
- Volunteer opportunities, charitable giving, social events, and continuous learning workshops.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →