Назад
Company hidden
3 часа назад

Member of Technical Staff, Kernels (AI)

200 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Kernels (AI): Designing and optimizing high-performance ML kernels and distributed compute infrastructure for large-scale diffusion language model training and inference with an accent on CUDA, CuTe, Triton, and low-precision arithmetic. Focus on reducing memory bandwidth bottlenecks, profiling GPU workloads, supporting distributed training, and maintaining reliable, scalable compute foundations.

Location: Bay Area, United States; in-office

Salary: $200,000–$350,000 annual base salary, plus equity and benefits

Company

hirify.global is a startup developing diffusion-based large language models, including Mercury, for fast and efficient language model inference and deployment.

What you will do

  • Design and implement custom ML kernels using CUDA, CuTe, and Triton for attention, matrix multiplication, gating, and normalization.
  • Develop compute primitives that reduce memory bandwidth bottlenecks and improve GPU kernel efficiency.
  • Improve infrastructure stability, scalability, reproducibility, and utilization across precision formats.
  • Support the distributed compute stack used for large-scale language model training and inference.

Requirements

  • BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
  • Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
  • Systems-level understanding of PyTorch or TensorFlow, plus experience profiling and optimizing ML systems.
  • Experience with low-precision formats such as FP8, INT8, or block floating point, or related compiler stacks including XLA or TVM.
  • Familiarity with data-parallel, model-parallel, and pipeline-parallel training, along with Python and C++, Rust, or Go.
  • Experience with Docker, Kubernetes, and CI/CD pipelines.

Nice to have

  • Experience building and maintaining language models with tens of billions of parameters or more.
  • Experience with distributed systems and AWS, GCP, or Azure.
  • Familiarity with PyTorch/XLA, DeepSpeed, or Megatron-LM.
  • Open-source contributions to PyTorch, DeepSpeed, XLA, or related deep learning infrastructure.

Culture & Benefits

  • Collaboration with leading AI researchers and inventors of diffusion models.
  • Opportunity to influence foundational AI technology and product direction.
  • Equity and competitive compensation in a rapidly growing startup.
  • Flexible vacation, paid time off, health, dental, and vision insurance.
  • 401(k) match, catered meals, commuter subsidies, and an inclusive culture.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →