Назад
Company hidden
6 дней назад

Machine Learning Systems Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Systems Engineer (AI): Building efficient runtime components and high-performance kernels that bring novel Core ML algorithms to Cerebras wafer-scale systems with an accent on compiler, runtime, communication, and low-level kernel optimization. Focus on translating research prototypes into robust implementations, profiling multi-layer performance bottlenecks, and optimizing large-scale training and low-latency inference.

Location: Hybrid in Sunnyvale, California, or Toronto, Canada

Company

hirify.global Systems develops wafer-scale AI hardware and software for high-speed machine learning training and inference.

What you will do

  • Design and implement runtime components and high-performance kernels for novel Core ML algorithms.
  • Translate research prototypes into efficient hirify.global implementations and GPU reference comparisons.
  • Profile and debug performance across ML frameworks, compilers, runtimes, communication layers, and kernels.
  • Optimize computation, memory movement, communication, and concurrency for large-scale training and low-latency inference.
  • Build benchmarks, instrumentation, automated tests, and numerical correctness validation.
  • Collaborate with researchers and compiler, runtime, kernel, and inference engineers on end-to-end capabilities and platform improvements.

Requirements

  • Bachelor’s, Master’s, PhD, or equivalent practical experience in computer science, computer engineering, electrical engineering, or a related field.
  • Experience developing high-performance systems software, ML systems, runtimes, compilers, or computational kernels.
  • Strong programming skills in C++ and Python.
  • Understanding of parallel programming, memory management, concurrency, data structures, and performance optimization.
  • Experience debugging and profiling complex software across multiple system layers.
  • Familiarity with modern machine learning architectures and frameworks such as PyTorch or JAX.

Nice to have

  • Experience with CUDA, Triton, low-level assembly, accelerator programming, or C-like domain-specific languages.
  • Experience with compiler internals, distributed runtimes, custom hardware interfaces, or HPC systems.
  • Knowledge of LLM training and inference, including attention, KV-cache management, parallel generation, or distributed execution.
  • Contributions to open-source systems, ML frameworks, compilers, or kernel libraries.

Culture & Benefits

  • Work on a wafer-scale AI platform designed beyond traditional GPU constraints.
  • Opportunities to publish and open-source AI research.
  • Access to one of the fastest AI supercomputers.
  • Job stability combined with startup vitality and a non-corporate work culture.
  • Commitment to an inclusive, diverse, and continuously learning workplace.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →