5 дней назад
Senior Machine Learning Systems Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Machine Learning Systems Engineer (AI/ML Platform): Building core infrastructure for creating, training, evaluating, deploying, and managing machine learning models and pipelines with an accent on scalable distributed systems, MLOps, and generative AI. Focus on designing fault-tolerant ML platforms, fine-tuning large language models, building retrieval-augmented generation systems, and optimizing performance across the ML lifecycle.
Location: Remote; listed locations include Sydney, Melbourne, and Brisbane, Australia, and Auckland, New Zealand. Hiring is possible in countries where has a legal entity.
Company
develops collaborative software products including Jira, Confluence, and Bitbucket, with a mission to help teams work more effectively.
What you will do
- Develop and refine core infrastructure for creating, training, evaluating, deploying, and managing machine learning models and pipelines.
- Collaborate with product teams such as Jira and Confluence to solve ML platform and infrastructure challenges.
- Curate ML datasets, fine-tune open-source large language models, and integrate proprietary LLMs.
- Lead projects from technical design through launch and contribute to company-wide engineering initiatives.
- Deliver code reviews, documentation, bug fixes, and technical mentorship.
Requirements
- 5+ years of experience building machine learning, AI infrastructure, platforms, or systems.
- Experience developing, deploying, and maintaining end-to-end ML systems, including data engineering, model serving, and monitoring.
- Extensive experience designing scalable, fault-tolerant, high-performance distributed systems for machine learning.
- Proficiency in Python and familiarity with PyTorch, TensorFlow, or JAX.
- Experience with MLOps, CI/CD pipelines, and automation for model training, deployment, and monitoring.
- Ability to diagnose and solve complex ML model and infrastructure performance problems.
Nice to have
- Experience with AWS, GCP, or Azure and AI/ML cloud services, including GPU compute.
- Experience with Spark, Ray, or Dask for large-scale data processing.
- Experience with generative AI frameworks, LLM fine-tuning, and retrieval-augmented generation systems.
- Experience with Go, Java, or Scala.
Culture & Benefits
- Flexible remote, office, or hybrid work options.
- Health and wellbeing resources.
- Paid volunteer days and community-focused benefits.
- Potential eligibility for benefits, bonuses, commissions, and equity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →