Назад
Company hidden
2 часа назад

Member of Technical Staff (ML Systems)

150 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff (ML Systems): Building and optimizing production inference systems for AI workloads with an accent on scheduling, memory management, runtime performance, and scalable model serving. Focus on designing execution strategies, improving KV cache efficiency, enabling new model architectures, and driving end-to-end performance across compiler, kernel, networking, and distributed systems.

Location: San Francisco, CA; on-site

Salary: $150K–$350K per year, plus equity

Company

hirify.global is building a multi-silicon neocloud for fast, efficient AI inference across heterogeneous hardware.

What you will do

  • Build and optimize production inference systems for latency, throughput, and efficiency.
  • Design execution strategies for batching, scheduling, concurrency, and resource utilization.
  • Improve KV cache management, memory efficiency, and runtime behavior at scale.
  • Enable new model architectures and inference techniques to run efficiently in production.
  • Partner with compiler, kernel, networking, and distributed systems engineers on end-to-end performance.
  • Influence the architecture of a large-scale AI infrastructure platform.

Requirements

  • Strong software engineering fundamentals.
  • Experience building or operating ML inference or model-serving systems.
  • Ability to reason about performance, memory usage, and system behavior under load.
  • Bachelor's degree in a relevant field, or equivalent education, training, and professional experience.
  • Software development experience in Python and C++.

Nice to have

  • Experience with TensorRT-LLM, vLLM, or custom serving systems.
  • Deep understanding of modern model architectures and attention mechanisms.
  • Experience with batching, scheduling, concurrency control, KV cache management, and memory placement.
  • Experience profiling and tuning latency- and throughput-critical systems.

Culture & Benefits

  • Early-stage environment with significant ownership over technical work.
  • Direct collaboration with a small group of highly capable colleagues.
  • Opportunity to work across domains and help shape the company's architecture and growth.
  • Equity offered as part of compensation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →