Назад
Company hidden
обновлено 17 дней назад

Staff/Senior Software Engineer, Machine Learning Platform (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Japan
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff/Senior Machine Learning Platform Engineer (Spark/Flink/Kubernetes): Building and scaling machine learning infrastructure for training, evaluation, deployment, and monitoring across batch and streaming pipelines with an accent on distributed data processing, job orchestration, and platform reliability. Focus on architecting high-throughput systems, operating ML workloads on Kubernetes, and developing observable, developer-friendly infrastructure at scale.

Location: Tokyo, Japan

Company

hirify.global is an AI-native Agentic AI as a Service company that develops intelligent software for business decision-making.

What you will do

  • Architect, implement, and scale Spark batch and Flink streaming pipelines processing billions of records daily for ML training and evaluation.
  • Design and operate job execution frameworks for model training, inference, and post-processing.
  • Build internal API servers and developer tools for orchestrating ML jobs on Kubernetes with Argo Workflows, Helm, and Terraform.
  • Design and monitor data infrastructure using ClickHouse and PostgreSQL.
  • Improve platform availability and observability with Prometheus and Grafana.
  • Collaborate with data scientists, product managers, and engineers; mentor junior engineers and promote LLM-based development tools.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field; a Master’s degree is preferred.
  • 4+ years of hands-on experience in data systems, machine learning infrastructure, or platform engineering.
  • Strong Python and/or Java skills with experience building large-scale production systems.
  • Practical experience with Spark, Flink, Kubernetes, GKE, Terraform, and Helm.
  • Experience with high-throughput data infrastructure such as ClickHouse or PostgreSQL.
  • Deep understanding of production ML pipelines, distributed job execution, architectural ownership, and cross-functional platform leadership.

Nice to have

  • Master’s degree in a relevant field.

Culture & Benefits

  • Opportunity to shape infrastructure supporting hundreds of ML models and billions of daily data records.
  • Use of LLM-based tools such as Claude Code, Codex, GitHub Copilot, and ChatGPT for development, documentation, and debugging.
  • Focus on engineering standards, ownership, scalability, and developer experience.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →