Назад
Company hidden
3 часа назад

Member of Technical Staff – ML Systems & Inference (AI)

250 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff – ML Systems & Inference (AI): Building production-grade ML inference and model serving systems for large-scale AI workloads with an accent on heterogeneous compute, scheduling, memory management, and runtime optimisation. Focus on reducing latency, improving throughput and resource utilisation, and enabling efficient production execution of new model architectures.

Location: San Francisco, CA — onsite

Salary: $250,000–$350,000 per year

Company

AI infrastructure company building inference systems and supporting production deployments for Fortune 500 and AI-native organisations.

What you will do

  • Design and build production-grade ML inference and model serving systems.
  • Optimise latency, throughput, and resource utilisation across large-scale AI workloads.
  • Develop batching, scheduling, concurrency, and runtime optimisation strategies.
  • Improve KV cache management, memory efficiency, and model execution behaviour.
  • Enable new model architectures and inference techniques to run efficiently in production.
  • Collaborate with compiler, kernel, networking, and distributed systems engineers on end-to-end performance.

Requirements

  • Strong software engineering fundamentals and significant ownership in a fast-moving environment.
  • Production experience building ML inference or model serving systems.
  • Deep understanding of system performance, memory behaviour, and optimisation under production workloads.
  • Experience with batching, scheduling, concurrency, KV cache management, and profiling latency- and throughput-critical systems.
  • Strong Python and C++ development experience.
  • Onsite work in San Francisco, California.

Nice to have

  • Experience with inference runtimes such as vLLM, TensorRT-LLM, or custom serving frameworks.

Culture & Benefits

  • Work in a small, highly technical engineering team.
  • Collaborate across compiler systems, GPU kernels, distributed scheduling, inference optimisation, and heterogeneous compute.
  • Build infrastructure used in production AI workloads at scale.
  • Join an exceptionally well-funded, early-stage AI infrastructure environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →