Назад
Company hidden
4 часа назад

Technical Director, Large-Scale AI Model Inferencing (AI)

219 000 - 351 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior/director
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Director, Large-Scale AI Model Inferencing (AI): Building full-stack memory solutions for production-scale AI inference across GPU HBM, host DRAM, CXL-attached pools, and NVMe/SSD tiers with an accent on model architecture, memory hierarchy, and inference performance. Focus on designing tiered memory systems, optimizing MoE and long-context workloads, developing analytical performance models, and leading technical strategy for production deployments.

Location: Daily onsite presence at the San Jose office/headquarters in San Jose, California, United States

Base pay range: $219,000–$351,000 USD per year, plus incentive opportunities.

Company

hirify.global develops technology solutions for smartphones, electric vehicles, hyperscale data centers, IoT devices, and other computing applications.

What you will do

  • Define full-stack AI memory requirements across GPU HBM, host DRAM, CXL-attached memory, and NVMe/SSD tiers.
  • Connect dense Transformer, MoE, SSM, hybrid, multimodal, and emerging model behavior to memory-system design.
  • Lead performance engineering for large-scale inference, including batching, KV-cache management, prefill/decode disaggregation, prefix caching, speculative decoding, and CUDA Graphs.
  • Develop benchmarks, simulators, analytical models, and production prototypes for latency, throughput, bandwidth, cache reuse, and cost-per-token optimization.
  • Design MoE expert-weight offloading and cache policies informed by routing behavior and real model workloads.
  • Set multi-year technical strategy, lead architecture reviews, contribute to inference open-source projects, and represent Samsung with customers and partners.

Requirements

  • BS in Computer, Electrical, or Electronic Engineering or Computer Science with 20 years of relevant experience; an MS with 18 years is preferred.
  • 12+ years in systems engineering and 4+ years of hands-on large-scale LLM inference or GPU systems performance experience.
  • First-principles understanding of Transformer internals, including KV-cache sizing, MHA/MQA/GQA/MLA tradeoffs, and activation-memory behavior.
  • Production experience with MoE behavior, expert parallelism, memory/storage systems, caching, tiering, paging, or storage engines.
  • Code-level expertise in at least one inference stack such as vLLM, SGLang, TensorRT-LLM, or llama.cpp, plus GPU/CPU profiling, NUMA, PCIe, and RDMA knowledge.
  • Strong technical communication skills, including executive-level narratives and customer-facing technical leadership.

Nice to have

  • Experience serving State Space Models or hybrid SSM-attention architectures.
  • Contributions to vLLM, SGLang, LMCache, HiCache, Mooncake, KTransformers, or llama.cpp.
  • Experience with CXL memory pooling, near-memory processing, SSD/NVMe cache tiers, or inference-grade QoS.
  • Background in memory or storage product companies delivering hardware-software co-designed solutions.

Culture & Benefits

  • Inclusive workplace focused on innovation, collaboration, curiosity, resilience, and diverse global perspectives.
  • Medical, dental, vision, and 401(k) benefits.
  • 4+ weeks of paid time off annually, holidays, and sick leave.
  • Family support, fertility and adoption assistance, medical travel support, emotional wellness resources, and virtual therapy.
  • Onsite café and gym, virtual fitness classes, charitable giving match, and a flexible work environment aligned with onsite requirements.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →