Назад
7 дней назад

Senior Machine Learning Engineer, Runtime and Serving Senior Machine Learning Engineer, Perception LLM/VLM (AI)

213 000 - 263 000$
Формат работы
onsite/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Machine Learning Engineer, Runtime and Serving Senior Machine Learning Engineer, Perception LLM/VLM (AI): Building high-performance ML runtime and serving systems for autonomous vehicle onboard compute and large-scale data centers with an accent on real-time inference, distributed serving, and hardware-aware optimization. Focus on extending JAX-native compilers and runtimes, optimizing C++/Python inference infrastructure, and solving latency, memory, concurrency, and tensor-management challenges.

Location: Mountain View, California, United States; onsite designation with a hybrid work arrangement also stated

Salary: $213,000–$263,000 USD annual base salary

Company

Waymo develops autonomous driving technology and operates a fully autonomous ride-hail service powered by the Waymo Driver.

What you will do

  • Architect and develop high-performance ML runtime and serving systems for autonomous vehicle compute and large-scale data centers.
  • Integrate inference runtimes while balancing real-time latency and memory constraints with high-throughput, concurrent offboard serving.
  • Drive migration of ML workloads to a JAX-native runtime architecture using OpenXLA/PjRT and TensorRT.
  • Collaborate with perception, planner, and research teams on system-level workloads and hardware-aware optimizations.
  • Build profiling and benchmarking tools to identify bottlenecks across the ML software stack.

Requirements

  • Bachelor’s or master’s degree in computer science, electrical engineering, deep learning, or a related field.
  • 5+ years of professional experience building, scaling, or maintaining ML systems and infrastructure.
  • 5+ years of production C++ programming experience.
  • 3+ years of production Python experience with major deep learning frameworks such as PyTorch or JAX.
  • Experience optimizing ML software for GPUs, TPUs, or custom silicon.
  • Experience building low-latency, highly concurrent distributed backend systems.

Nice to have

  • PhD in computer science, electrical engineering, deep learning, or a related field.
  • Experience modifying ML compilers, runtimes, or inference engines such as TensorRT, ONNX Runtime, OpenXLA/PjRT, or TVM.
  • Experience building or scaling LLM serving systems, distributed inference, KV or prefix caching, and continuous batching.
  • Experience with custom kernel development using CUDA, CUDA Tile, Triton, JAX, or Pallas.
  • Experience architecting unified serving APIs and optimizing tensor buffer management for multi-model inference pipelines.

Culture & Benefits

  • Work across the ML stack with teams in perception, planning, research, and simulation.
  • Eligibility for a discretionary annual bonus program.
  • Eligibility for an equity incentive plan.
  • Generous company benefits program, subject to eligibility requirements.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →