3 часа назад
Member of Technical Staff, Kernels (AI)
200 000 - 350 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, Kernels (AI): Designing and optimizing high-performance ML kernels and distributed compute infrastructure for large-scale diffusion language model training and inference with an accent on CUDA, CuTe, Triton, and low-precision arithmetic. Focus on reducing memory bandwidth bottlenecks, profiling GPU workloads, supporting distributed training, and maintaining reliable, scalable compute foundations.
Location: Bay Area, United States; in-office
Salary: $200,000–$350,000 annual base salary, plus equity and benefits
Company
is a startup developing diffusion-based large language models, including Mercury, for fast and efficient language model inference and deployment.
What you will do
- Design and implement custom ML kernels using CUDA, CuTe, and Triton for attention, matrix multiplication, gating, and normalization.
- Develop compute primitives that reduce memory bandwidth bottlenecks and improve GPU kernel efficiency.
- Improve infrastructure stability, scalability, reproducibility, and utilization across precision formats.
- Support the distributed compute stack used for large-scale language model training and inference.
Requirements
- BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
- Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
- Systems-level understanding of PyTorch or TensorFlow, plus experience profiling and optimizing ML systems.
- Experience with low-precision formats such as FP8, INT8, or block floating point, or related compiler stacks including XLA or TVM.
- Familiarity with data-parallel, model-parallel, and pipeline-parallel training, along with Python and C++, Rust, or Go.
- Experience with Docker, Kubernetes, and CI/CD pipelines.
Nice to have
- Experience building and maintaining language models with tens of billions of parameters or more.
- Experience with distributed systems and AWS, GCP, or Azure.
- Familiarity with PyTorch/XLA, DeepSpeed, or Megatron-LM.
- Open-source contributions to PyTorch, DeepSpeed, XLA, or related deep learning infrastructure.
Culture & Benefits
- Collaboration with leading AI researchers and inventors of diffusion models.
- Opportunity to influence foundational AI technology and product direction.
- Equity and competitive compensation in a rapidly growing startup.
- Flexible vacation, paid time off, health, dental, and vision insurance.
- 401(k) match, catered meals, commuter subsidies, and an inclusive culture.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →