Назад
Company hidden
9 дней назад

Principal Machine Learning Engineer (On-Device AI)

250 300 - 325 400$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Machine Learning Engineer (On-Device AI): Building and optimizing on-device generative AI inference systems for browser-native game experiences with an accent on WebGPU deployment, model compression, GPU kernel tuning, and real-time engine integration. Focus on reducing latency, memory, and power consumption across mobile and desktop hardware, designing runtime scheduling and zero-copy pipelines, and productionizing transformer and diffusion models.

Location: Mountain View, California, USA

Base salary: $250,300–$325,400 in Zone A; $222,600–$289,400 in Zone B; $197,400–$256,600 in Zone C. The range reflects annual base salary.

Company

hirify.global develops a leading game engine used to create games, XR and web experiences, and 3D applications across industries including automotive, manufacturing, and healthcare.

What you will do

  • Own the end-to-end optimization pipeline for transformer and diffusion models, from export and graph transformation through operator fusion, quantization, pruning, and hardware-specific kernel tuning.
  • Develop and optimize WebGPU compute shaders and native GPU kernels across mobile NPUs, mobile GPUs, and desktop GPUs using profiling and frame-capture tools.
  • Design the integration between ML inference runtimes and the game engine, including scheduling, threading, memory pooling, zero-copy buffers, and frame-budget management.
  • Build model packaging, asset pipelines, device fallbacks, SKU-aware capability tiers, telemetry, and automated on-device benchmarking.
  • Partner with research scientists to productionize new architectures and assess efficient inference methods against latency, quality, memory, and power targets.
  • Lead and mentor engineers while establishing performance standards, benchmarking practices, code-review guidelines, and regression gates.

Requirements

  • 8+ years of software or ML engineering experience, including at least 4 years in on-device, edge-inference, or real-time performance-critical systems.
  • Production experience deploying transformer or diffusion models on mobile, desktop, or embedded hardware.
  • Hands-on WebGPU experience, including ONNX Runtime Web, Transformers.js, WebLLM, or TensorFlow.js, plus WGSL shader development; equivalent native GPU or compute API expertise may be considered.
  • Deep experience with an inference runtime such as ONNX Runtime, CoreML, TFLite, or ExecuTorch, including operator fusion, memory layout, and runtime scheduling.
  • Strong knowledge of GPU or compute APIs, model optimization, target hardware, TypeScript/JavaScript, WGSL, and Python.
  • Professional English proficiency is required for regular communication with colleagues and partners worldwide.

Nice to have

  • Experience with neural rendering, world models, real-time generative pipelines, or on-device diffusion.
  • Game-engine or real-time graphics experience with hirify.global, Unreal, or custom engines.
  • Contributions to ML inference, runtime, GPU, or WebGPU open-source projects.
  • Familiarity with WebGPU compute features, compiler stacks such as MLIR, TVM, IREE, or XLA, and device-farm benchmarking infrastructure.

Culture & Benefits

  • Health, life, and disability insurance, with offerings varying by country and employment status.
  • Employee stock ownership and retirement or pension plans.
  • Vacation, personal days, parental leave, and family-care programs.
  • Mental health support, employee resource groups, training and development programs, and an employee assistance program.
  • Commute subsidy, office snacks, and volunteering and donation matching programs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →