2 дня назад
ML Operations Engineer (AI/LLM)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Operations Engineer (AI/LLM) (Cloud-Native ML Serving): Building and operating production model-serving and data-orchestration platforms for AI and LLM features serving tens of millions of users with an accent on deployment automation, inference performance, and reliability. Focus on optimizing quantization, dynamic batching, and model evaluation while designing scalable Kubernetes and Terraform infrastructure for emerging agentic AI workloads.
Location: Minato City, Tokyo, Japan; Workplace: Hybrid; Office: Roppongi
Company
operates a marketplace platform focused on circulating value and creating opportunities through technology.
What you will do
- Own end-to-end orchestration of model inference, including data retrieval through BigQuery, Bigtable, and Valkey.
- Operate production model serving on NVIDIA and TPU cloud-native stacks using Triton Inference Server, TensorRT-LLM, and JAX/TPU Gateways.
- Build CI/CD, rollout, rollback, provisioning, and lifecycle-management workflows with Terraform and Kubernetes.
- Optimize inference latency, throughput, and cost through model compilation, quantization, dynamic batching, and performance regression detection.
- Build monitoring, alerting, SLOs, on-call processes, incident response, and model-quality evaluation for serving systems.
- Partner with ML engineers and researchers to bring experimental models into reliable production services.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 5+ years of software engineering experience, including production MLOps covering model deployment, serving, and CI/CD in cloud environments.
- Experience designing and operating large-scale, highly available distributed systems, including observability, SLOs, and incident response.
- Strong experience with Kubernetes, Docker, Python, and Terraform.
- Excellent written and verbal communication skills.
- English proficiency at CEFR B2 is required.
Nice to have
- Experience integrating ML serving with large-scale data warehouses, wide-column stores, or in-memory caches.
- Experience with TensorRT-LLM, quantization, JAX, model-inference gateways, and orchestrators.
- 2+ years of operating production GenAI or LLM workloads, including token-throughput and cost optimization.
- Experience with LLM evaluation, guardrails, quality monitoring, RAG, vector search, or agentic AI workloads.
- Experience partnering with research or data science teams; a Master’s degree or Ph.D. is a plus.
Culture & Benefits
- Full flextime with no core working hours.
- Engineering culture built around product focus, continuous growth, mechanism-based problem solving, and open collaboration.
- Work on AI and LLM infrastructure supporting products used by tens of millions of users.
Hiring process
- Application screening.
- Engineering skill assessment through HackerRank or GitHub.
- Interviews, reference check, and offer decision.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 часов назад
Software Engineer 2 (AI)
103 000 - 155 000CAD
6 часов назад
Software Engineer 3 (Enterprise AI)
132 000 - 198 000CAD
5 дней назад
AI Research Scientist/Engineer
5 дней назад
Revenue Operations Engineer (AI)
90 000 - 115 000$
5 дней назад
GTM Systems Engineer (AI)
90 000 - 115 000$
Databricks
3 дня назад