Назад
Company hidden
12 часов назад

Solutions Architect (AI Infrastructure)

Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US/SK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Solutions Architect (AI Infrastructure): Enabling US customers to deploy AI and LLM models on FuriosaAI's RNGD NPU using the Furiosa SDK with an accent on proof-of-concepts, benchmarking, debugging, and customer integration. Focus on optimizing inference stacks and agent frameworks, translating accelerator capabilities into business value, and feeding customer requirements back to product and engineering teams.

Location: Santa Clara, California; authorized to work in the US; travel to customer sites and Seoul HQ required periodically

Company

hirify.global develops RNGD AI chips and servers, along with software for deploying and optimizing AI and LLM inference workloads.

What you will do

  • Own end-to-end technical enablement for US customers deploying AI models on hirify.global's RNGD NPU through the Furiosa SDK.
  • Develop proof-of-concepts, benchmarking studies, and live debugging sessions in customer environments.
  • Support the US business development and sales organization during pre-sales and enterprise evaluations.
  • Translate technical capabilities into business value for ML engineers, engineering leaders, and C-suite audiences.
  • Train customers on integration patterns, optimization workflows, and post-purchase best practices.
  • Provide technical feedback from US customers to product and engineering teams at Seoul HQ.

Requirements

  • 2–5 years of experience in a US customer-facing technical role such as Solutions Architect, Sales Engineer, or Forward Deployed Engineer at an AI infrastructure, cloud, or semiconductor company.
  • Current knowledge of the AI and LLM landscape, including model releases, inference frameworks, and serving stacks.
  • Hands-on experience with vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar inference technologies.
  • Experience with agent and orchestration frameworks such as LangChain, LlamaIndex, LangGraph, AutoGen, or MCP-based tooling.
  • Proficiency in Python and familiarity with PyTorch and TensorFlow.
  • US work authorization and willingness to travel to customer sites and Seoul HQ periodically, plus strong written and verbal communication skills.

Nice to have

  • Experience at a US AI chip company, cloud silicon team, or AI infrastructure startup.
  • Familiarity with NPU/GPU accelerator ecosystems, PCIe integration, and data center hardware deployment.
  • Experience with inference optimization, including quantization, kernel tuning, batching, and memory bandwidth optimization.
  • Proficiency in C, C++, or Rust.
  • Experience working with distributed or cross-time-zone engineering teams.

Culture & Benefits

  • Work directly with AI infrastructure hardware and software in real-world customer environments.
  • Engage with ML engineers, frontier AI labs, enterprise customers, and executive stakeholders.
  • Participate in US technical forums, AI conferences, and customer workshops.
  • Collaborate with product and engineering teams across the US and Seoul.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →