3 дня назад
Senior Research Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Research Engineer (AI): Adapting and optimizing language and vision models for efficient inference on Cerebras AI hardware with an accent on speculative decoding, model pruning and compression, sparse attention, and sparsity-driven techniques. Focus on designing inference algorithms, profiling model performance, and building low-latency, high-throughput systems for large-scale workloads.
Location: Hybrid role in Toronto, ON, Canada, or Sunnyvale, CA, USA
Company
Systems develops large-scale AI accelerator hardware and systems designed to deliver high-speed model training and inference.
What you will do
- Adapt, design, implement, and optimize transformer architectures for NLP and computer vision on hardware.
- Research and prototype inference algorithms and model architectures focused on speculative decoding, pruning and compression, sparse attention, and sparsity.
- Train models to convergence, run hyperparameter sweeps, and analyze experimental results.
- Bring up new models, validate functional correctness, and troubleshoot integration issues on the system.
- Profile and optimize model code to maximize throughput and minimize inference latency.
- Develop diagnostic tools and collaborate with software, hardware, and product teams to deliver inference projects.
Requirements
- Relevant bachelor's degree with 7+ years of ML software development experience, master's degree with 4+ years of software development experience, PhD with 2+ years of relevant experience, or equivalent practical experience.
- 4+ years of experience testing, maintaining, or launching software products, including 2+ years in software design and architecture.
- 3+ years of machine-learning-focused software development experience, including deep learning, large language models, or computer vision.
- Strong programming skills in Python and/or C++, experience with generative AI and ML systems, and proficiency in PyTorch, Transformers, vLLM, or SGLang.
- Deep understanding of transformer-based models, inference optimization, specialized hardware performance, sparse attention, pruning and compression, and speculative decoding.
- Evidence of ML research impact through publications, open-source contributions, or high-quality preprints.
Nice to have
- Experience with large language models, mixture-of-experts models, multimodal learning, or AI agents.
- Experience with quantization, post-training techniques, inference evaluations, and large-scale model deployment.
- Triton or CUDA experience.
- Experience independently taking complex ML or inference projects from prototype to production-quality implementation.
Culture & Benefits
- Opportunity to build AI infrastructure beyond traditional GPU constraints.
- Support for publishing and open-sourcing AI research.
- Work with a high-performance AI supercomputer platform.
- Startup vitality combined with job stability.
- Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Anthropic
8 дней назад
Research Engineer (AI)
350 000 - 850 000$
Baseten
8 дней назад
Software Engineer (AI)
180 000 - 360 000$
8 дней назад
AI Systems Researcher/Engineer (ML)
6 дней назад
Applied Research Engineer (AI)
Anthropic
6 дней назад
Research Engineer / Research Scientist, RL Frontiers (AI)
500 000 - 850 000$
6 дней назад
Applied Robotics Engineer (AI)
180 000 - 300 000$