Назад
Company hidden
6 часов назад

Senior Edge Inference Engineer (AI)

Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Edge Inference Engineer (AI/C++): Implementing and optimizing machine-learning inference kernels and pipelines for CPU, NPU, and GPU architectures on resource-constrained edge devices with an accent on quantization, memory efficiency, and hardware-level optimization. Focus on achieving sub-100ms time-to-first-token, supporting new model architectures, and upstreaming production-grade improvements to llama.cpp.

Location: Hybrid in San Francisco or Boston; remote work is available, and other locations may be considered

Company

hirify.global, spun out of MIT CSAIL, builds general-purpose AI systems optimized for data centers and on-device hardware.

What you will do

  • Implement and optimize inference kernels for CPU, NPU, and GPU architectures across diverse edge devices.
  • Develop INT4, INT8, and FP8 quantization strategies that reduce memory usage while preserving model quality.
  • Contribute to llama.cpp and other open-source inference frameworks, including support for audio and vision model architectures.
  • Profile and optimize end-to-end inference pipelines to achieve sub-100ms time-to-first-token on target devices.
  • Collaborate with ML researchers to identify optimization opportunities for Liquid Foundation Models.
  • Own major optimization workstreams from design through production deployment and upstream contribution.

Requirements

  • 5+ years of systems programming experience with strong C++ proficiency.
  • Experience with embedded software or resource-constrained systems.
  • Understanding of ML fundamentals, including matrix operations, attention mechanisms, and quantization.
  • Knowledge of cache hierarchies, memory bandwidth, SIMD, and vectorization.
  • Ability to work autonomously, diagnose performance bottlenecks, and ship maintainable production code.
  • Experience delivering measurable latency or memory improvements on an edge device class.

Nice to have

  • Contributions to llama.cpp, ExecuTorch, or similar inference frameworks.
  • Rust systems programming experience.
  • Experience with custom TPU or NPU accelerator development.
  • A quantitative degree in mathematics, physics, or a related field combined with engineering experience.

Culture & Benefits

  • High-ownership work involving novel model architectures and optimization strategies.
  • Production code runs on real customer devices and contributes to open-source projects.
  • Competitive base salary with equity in a unicorn-stage company.
  • 100% coverage of medical, dental, and vision premiums for employees and dependents.
  • 401(k) matching up to 4% of base pay.
  • Unlimited PTO and company-wide Refill Days.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →