Назад
Company hidden
21 день назад

Forward Deployment Engineer (AI)

10 500 000 - 13 000 000JPY
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Japan
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Forward Deployment Engineer (AI): Designing and deploying production generative AI applications on SambaNova's SN40L platform and SambaStack with an accent on LLM orchestration, RAG pipelines, agentic systems, and inference optimization. Focus on benchmarking model performance, troubleshooting across model, software, and hardware layers, and translating customer requirements into scalable architectures and product feedback.

Location: Tokyo Prefecture, Japan; willingness to travel up to 50% to customer sites is required.

Salary: ¥10,500,000–¥13,000,000 JPY per year

Company

hirify.global develops full-stack AI inference infrastructure based on RDU hardware, SambaRack systems, and the SambaStack enterprise inference serving platform.

What you will do

  • Embed with strategic enterprise customers to design, build, and deploy production generative AI applications on the SN40L platform and SambaStack.
  • Architect LLM-powered workflows, including RAG pipelines, multi-agent systems, fine-tuning workflows, and coding solutions.
  • Optimize inference performance by benchmarking throughput, latency, and accuracy against customer requirements and competitor baselines.
  • Troubleshoot production issues across model, software, and hardware layers as the primary technical escalation point.
  • Translate customer needs into product requirements and engineering feedback while partnering with sales and solutions engineering teams.
  • Develop reusable accelerators, reference architectures, and playbooks, and present technical findings to customers and internal teams.

Requirements

  • 5+ years of hands-on engineering experience shipping production AI/ML systems.
  • Deep expertise in LLM orchestration, RAG, agentic frameworks, prompt engineering, and evaluation pipelines.
  • Strong knowledge of model training, fine-tuning, inference optimization, quantization, and performance benchmarking.
  • Proficiency in Python; working knowledge of C++ or CUDA is beneficial for hardware-layer debugging.
  • Experience with AWS, Azure, or GCP, plus containerization, Kubernetes, Docker, and MLOps tooling.
  • Bachelor's or graduate degree in Computer Science, Electrical Engineering, Mathematics, Physics, or equivalent practical experience.

Nice to have

  • Experience with AI accelerators or custom silicon, including TPUs.
  • CUDA or low-level GPU programming experience.
  • Familiarity with vLLM or SGLang.
  • Enterprise AI deployment experience in regulated industries.

Culture & Benefits

  • Work on an integrated hardware and software platform for generative AI inference and high-performance computing.
  • Collaborate with enterprise customers, Product and Engineering teams, Account Executives, and Solutions Engineers.
  • Represent the platform at customer briefings, industry conferences, and technical events.
  • Benefits described in the posting apply to US-based full-time employment positions and include medical, dental, vision, disability, life, HSA, FSA, wellness, and counseling programs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →