3 дня назад
GPU Kernel Engineer (AI)
65$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
GPU Kernel Engineer (AI) (CUDA/Triton): Reviewing, debugging, and evaluating high-performance GPU and accelerator kernels for AI workloads with an accent on numerical correctness, memory optimization, and benchmarking. Focus on translating kernels between frameworks, migrating implementations across hardware, profiling bottlenecks, and validating compilation, runtime behavior, and performance targets.
Location: Remote from Argentina, Brazil, Chile, Colombia, Ecuador, Mexico, Portugal, Spain, or Uruguay
Compensation: $65 per hour
Company
is running a specialized project focused on evaluating and improving high-performance GPU and accelerator kernels for AI workloads.
What you will do
- Review, implement, debug, and evaluate GPU and accelerator kernels for correctness and performance.
- Optimize CUDA and Triton kernels through profiling, operator fusion, memory hierarchy improvements, and hardware-aware tuning.
- Translate kernels between frameworks and migrate implementations across NVIDIA GPUs and custom accelerator platforms.
- Compare outputs with reference implementations and verify numerical tolerance thresholds.
- Investigate compilation, driver, memory, shape, and runtime issues.
- Assess benchmarks, identify bottlenecks, and provide clear technical feedback on optimization opportunities and task difficulty.
Requirements
- 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels.
- Strong experience with at least two of CUDA, Triton, NKI/AWS Neuron, and Pallas/JAX.
- Strong understanding of GPU performance optimization, including memory bandwidth, compute throughput, occupancy, shared memory, register pressure, memory coalescing, and bank conflicts.
- Experience with Nsight, NCU, roofline analysis, or framework-native profiling tools.
- Strong understanding of floating-point numerical correctness and tolerance thresholds.
- Experience debugging kernel compilation and runtime issues and distinguishing software defects from environment problems and optimization challenges.
Nice to have
- Experience with AWS Trainium, TPU, JAX, or other accelerator ecosystems.
- Compiler engineering experience and familiarity with MLIR, XLA, or intermediate representation lowering.
- Contributions to GPU or ML kernel libraries, including cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls.
- Experience with AI model evaluation, RLHF, or technical benchmark development.
Culture & Benefits
- Remote, part-time, project-based consulting engagement.
- Work focused on GPU kernels, performance engineering, debugging, and technical evaluation.
- Close-to-hardware work aimed at pushing AI compute systems toward their performance limits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →