Назад
Company hidden
2 дня назад

ML Operations Engineer (AI/LLM)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Japan
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Operations Engineer (AI/LLM) (Cloud-Native ML Serving): Building and operating production model-serving and data-orchestration platforms for AI and LLM features serving tens of millions of users with an accent on deployment automation, inference performance, and reliability. Focus on optimizing quantization, dynamic batching, and model evaluation while designing scalable Kubernetes and Terraform infrastructure for emerging agentic AI workloads.

Location: Minato City, Tokyo, Japan; Workplace: Hybrid; Office: Roppongi

Company

hirify.global operates a marketplace platform focused on circulating value and creating opportunities through technology.

What you will do

  • Own end-to-end orchestration of model inference, including data retrieval through BigQuery, Bigtable, and Valkey.
  • Operate production model serving on NVIDIA and TPU cloud-native stacks using Triton Inference Server, TensorRT-LLM, and JAX/TPU Gateways.
  • Build CI/CD, rollout, rollback, provisioning, and lifecycle-management workflows with Terraform and Kubernetes.
  • Optimize inference latency, throughput, and cost through model compilation, quantization, dynamic batching, and performance regression detection.
  • Build monitoring, alerting, SLOs, on-call processes, incident response, and model-quality evaluation for serving systems.
  • Partner with ML engineers and researchers to bring experimental models into reliable production services.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 5+ years of software engineering experience, including production MLOps covering model deployment, serving, and CI/CD in cloud environments.
  • Experience designing and operating large-scale, highly available distributed systems, including observability, SLOs, and incident response.
  • Strong experience with Kubernetes, Docker, Python, and Terraform.
  • Excellent written and verbal communication skills.
  • English proficiency at CEFR B2 is required.

Nice to have

  • Experience integrating ML serving with large-scale data warehouses, wide-column stores, or in-memory caches.
  • Experience with TensorRT-LLM, quantization, JAX, model-inference gateways, and orchestrators.
  • 2+ years of operating production GenAI or LLM workloads, including token-throughput and cost optimization.
  • Experience with LLM evaluation, guardrails, quality monitoring, RAG, vector search, or agentic AI workloads.
  • Experience partnering with research or data science teams; a Master’s degree or Ph.D. is a plus.

Culture & Benefits

  • Full flextime with no core working hours.
  • Engineering culture built around product focus, continuous growth, mechanism-based problem solving, and open collaboration.
  • Work on AI and LLM infrastructure supporting products used by tens of millions of users.

Hiring process

  • Application screening.
  • Engineering skill assessment through HackerRank or GitHub.
  • Interviews, reference check, and offer decision.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →