Назад
Company hidden
3 дня назад

Senior Research Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Research Engineer (AI): Adapting and optimizing language and vision models for efficient inference on Cerebras AI hardware with an accent on speculative decoding, model pruning and compression, sparse attention, and sparsity-driven techniques. Focus on designing inference algorithms, profiling model performance, and building low-latency, high-throughput systems for large-scale workloads.

Location: Hybrid role in Toronto, ON, Canada, or Sunnyvale, CA, USA

Company

hirify.global Systems develops large-scale AI accelerator hardware and systems designed to deliver high-speed model training and inference.

What you will do

  • Adapt, design, implement, and optimize transformer architectures for NLP and computer vision on hirify.global hardware.
  • Research and prototype inference algorithms and model architectures focused on speculative decoding, pruning and compression, sparse attention, and sparsity.
  • Train models to convergence, run hyperparameter sweeps, and analyze experimental results.
  • Bring up new models, validate functional correctness, and troubleshoot integration issues on the hirify.global system.
  • Profile and optimize model code to maximize throughput and minimize inference latency.
  • Develop diagnostic tools and collaborate with software, hardware, and product teams to deliver inference projects.

Requirements

  • Relevant bachelor's degree with 7+ years of ML software development experience, master's degree with 4+ years of software development experience, PhD with 2+ years of relevant experience, or equivalent practical experience.
  • 4+ years of experience testing, maintaining, or launching software products, including 2+ years in software design and architecture.
  • 3+ years of machine-learning-focused software development experience, including deep learning, large language models, or computer vision.
  • Strong programming skills in Python and/or C++, experience with generative AI and ML systems, and proficiency in PyTorch, Transformers, vLLM, or SGLang.
  • Deep understanding of transformer-based models, inference optimization, specialized hardware performance, sparse attention, pruning and compression, and speculative decoding.
  • Evidence of ML research impact through publications, open-source contributions, or high-quality preprints.

Nice to have

  • Experience with large language models, mixture-of-experts models, multimodal learning, or AI agents.
  • Experience with quantization, post-training techniques, inference evaluations, and large-scale model deployment.
  • Triton or CUDA experience.
  • Experience independently taking complex ML or inference projects from prototype to production-quality implementation.

Culture & Benefits

  • Opportunity to build AI infrastructure beyond traditional GPU constraints.
  • Support for publishing and open-sourcing AI research.
  • Work with a high-performance AI supercomputer platform.
  • Startup vitality combined with job stability.
  • Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →