обновлено 17 дней назад
Staff/Senior Software Engineer, Machine Learning Platform (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff/Senior Machine Learning Platform Engineer (Spark/Flink/Kubernetes): Building and scaling machine learning infrastructure for training, evaluation, deployment, and monitoring across batch and streaming pipelines with an accent on distributed data processing, job orchestration, and platform reliability. Focus on architecting high-throughput systems, operating ML workloads on Kubernetes, and developing observable, developer-friendly infrastructure at scale.
Location: Tokyo, Japan
Company
is an AI-native Agentic AI as a Service company that develops intelligent software for business decision-making.
What you will do
- Architect, implement, and scale Spark batch and Flink streaming pipelines processing billions of records daily for ML training and evaluation.
- Design and operate job execution frameworks for model training, inference, and post-processing.
- Build internal API servers and developer tools for orchestrating ML jobs on Kubernetes with Argo Workflows, Helm, and Terraform.
- Design and monitor data infrastructure using ClickHouse and PostgreSQL.
- Improve platform availability and observability with Prometheus and Grafana.
- Collaborate with data scientists, product managers, and engineers; mentor junior engineers and promote LLM-based development tools.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field; a Master’s degree is preferred.
- 4+ years of hands-on experience in data systems, machine learning infrastructure, or platform engineering.
- Strong Python and/or Java skills with experience building large-scale production systems.
- Practical experience with Spark, Flink, Kubernetes, GKE, Terraform, and Helm.
- Experience with high-throughput data infrastructure such as ClickHouse or PostgreSQL.
- Deep understanding of production ML pipelines, distributed job execution, architectural ownership, and cross-functional platform leadership.
Nice to have
- Master’s degree in a relevant field.
Culture & Benefits
- Opportunity to shape infrastructure supporting hundreds of ML models and billions of daily data records.
- Use of LLM-based tools such as Claude Code, Codex, GitHub Copilot, and ChatGPT for development, documentation, and debugging.
- Focus on engineering standards, ownership, scalability, and developer experience.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →