Назад
Company hidden
6 дней назад

Machine Learning Engineer (Model Quantization)

160 000 - 215 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer (Model Quantization): Developing hardware-aware post-training quantization and model adaptation methods for LLMs, diffusion models, and other AI workloads on Neurophos optical inference engines with an accent on numerical optimization, low-precision compute, and model deployment. Focus on designing reproducible experiments, minimizing accuracy loss during precision reduction, optimizing GEMM operations, and co-optimizing models with photonic hardware.

Location: Austin, TX or Sunnyvale, CA; full-time onsite position

Salary: $160K–$215K annually, plus equity

Company

hirify.global is an AI hardware startup developing silicon photonics and programmable metasurface-based optical inference engines for highly efficient, high-throughput AI computation.

What you will do

  • Develop hardware-aware post-training quantization methods for LLMs, diffusion models, and other machine learning applications.
  • Investigate non-convex, discrete, constrained, and second-order optimization approaches for quantization and preconditioning.
  • Design controlled numerical experiments and build research-quality, reproducible implementation and evaluation harnesses.
  • Adapt open-source and customer models across PyTorch, Triton, JAX, and other frameworks.
  • Develop re-quantization, retraining, and model adaptation techniques that minimize accuracy loss during precision reduction.
  • Collaborate with hardware, software, and architecture teams to optimize GEMM operations and model architectures for optical compute.

Requirements

  • PhD or equivalent research experience in machine learning, applied mathematics, optimization, numerical analysis, computer science, or a related field.
  • 5+ years of machine learning engineering experience, including at least 3 years focused on model optimization and deployment.
  • Research or advanced engineering experience in neural network quantization, model compression, numerical optimization, or efficient inference.
  • Strong knowledge of numerical linear algebra and experience with advanced optimization methods.
  • Strong proficiency in PyTorch and familiarity with JAX, Triton, and TensorFlow.
  • Hands-on experience with transformer architectures, LLMs, diffusion models, controlled numerical experiments, and research collaboration.

Nice to have

  • Experience with INT8, FP8, or lower-precision inference optimization.
  • Background in analog or optical computing, in-memory computing, or matrix-vector multiplication acceleration.
  • Knowledge of randomized numerical linear algebra, sketching, structured transforms, vector quantization, lattice methods, learned codebooks, or rate-distortion techniques.
  • Publications in quantization, optimization, numerical linear algebra, model compression, or efficient machine learning.
  • Experience with large-scale batch inference and LLM prefill versus decode optimization.

Culture & Benefits

  • Health plan premiums covered at 100% for employees and dependents, with HSA contributions.
  • Unlimited paid time off.
  • 401(k) matching and stock option opportunities.
  • Dental, vision, life, hospital, critical illness, and accident insurance options.
  • Flexible benefits selection with cash-back options for unused plans.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →