3 дня назад
Senior MLOps Engineer (AI/ML)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior MLOps Engineer (AI/ML): Building and operating end-to-end ML training, deployment, serving, and monitoring infrastructure for multimodal spatial frontier models with an accent on distributed GPU workloads, geospatial data, and reliable production delivery. Focus on automating CI/CD and reproducible environments, scaling inference across customers and regions, and solving performance, data quality, observability, and deployment challenges.
Location: Hybrid in Sydney, Australia
Company
develops machine-learning-powered, physics-enabled digital twins of electricity grids to help utilities assess infrastructure risks, optimise investments, and improve climate resilience.
What you will do
- Build and operate end-to-end ML training, evaluation, and deployment pipelines from data ingestion to customer environments.
- Automate model production workflows with CI/CD, model and artefact registries, reproducible environments, and infrastructure as code.
- Implement monitoring, alerting, drift detection, and data-quality checks for production models, responding to operational failures.
- Manage distributed training infrastructure, GPU clusters, scheduling, and resource utilisation across cloud, on-premises, and neocloud environments.
- Deploy and scale inference services across regions and customers, balancing latency, cost, and data residency requirements.
- Improve ML engineering tools, workflows, and documentation as the team scales.
Requirements
- Hands-on experience operating ML training pipelines, model-serving systems, and monitoring in production.
- Strong Python skills and working knowledge of PyTorch or an equivalent framework.
- Experience with distributed training, GPU infrastructure, scheduling, resource management, and troubleshooting throughput or memory issues.
- Experience with AWS, GCP, or Azure; Kubernetes, Docker, and infrastructure as code.
- Experience with production model monitoring, data-quality frameworks, and preparing training data for ML readiness.
- Strong engineering judgement, maintainable coding practices, and awareness of failure modes.
Nice to have
- CUDA or kernel-level optimisation experience.
- Exposure to point-cloud or geospatial data.
- Experience deploying into regulated or air-gapped customer environments.
Culture & Benefits
- Competitive salary.
- Meaningful ESOP.
- Fully flexible work environment with an office in Redfern.
- Regular office events.
- Opportunity to work on a complex product supporting critical infrastructure and climate resilience.
Hiring process
- Apply through the online application process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →