Назад
Company hidden
6 дней назад

Senior Machine Learning Engineer (AI)

250 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Machine Learning Engineer (AI): Building production-grade ML inference and model serving infrastructure for heterogeneous compute with an accent on latency, throughput, scheduling, memory management, and runtime optimisation. Focus on designing scalable inference systems, improving KV cache efficiency, enabling new model architectures, and coordinating compiler, kernel, networking, and distributed systems work.

Location: San Francisco, California, United States — onsite

Salary: $250,000–$350,000 per year

Company

AI infrastructure company building production inference systems for large-scale AI workloads and heterogeneous compute.

What you will do

  • Design and build production-grade ML inference and model serving systems.
  • Optimise latency, throughput, resource utilisation, and memory efficiency across large-scale AI workloads.
  • Develop batching, scheduling, concurrency, runtime optimisation, and KV cache management strategies.
  • Enable new model architectures and inference techniques to run efficiently in production.
  • Collaborate with compiler, kernel, networking, and distributed systems engineers on end-to-end performance.
  • Help shape the architecture of a next-generation AI inference platform.

Requirements

  • Strong software engineering fundamentals and significant ownership in a fast-moving environment.
  • Production experience building ML inference or model serving systems.
  • Deep understanding of system performance, memory behaviour, and production optimisation.
  • Experience with batching, scheduling, concurrency, KV cache management, and profiling latency- and throughput-critical systems.
  • Strong Python and C++ development experience.

Nice to have

  • Experience with vLLM, TensorRT-LLM, or custom inference serving frameworks.
  • Experience working with compiler systems, GPU kernels, distributed scheduling, or heterogeneous compute.

Culture & Benefits

  • Small, highly technical, world-class engineering team.
  • Early-stage environment with substantial technical ownership.
  • Work on production deployments for Fortune 500 and AI-native organisations.
  • Opportunity to solve challenging AI infrastructure and high-performance systems problems.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →