обновлено 10 дней назад
Parallel Computing Engineer (CUDA)
130 000 - 180 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Parallel Computing Engineer (CUDA/HPC): Designing and optimizing GPU kernels, distributed training and inference architectures, and large-scale AI and scientific computing workloads with an accent on CUDA performance, GPU memory management, and multi-GPU scaling. Focus on profiling with NVIDIA Nsight tools, integrating custom operators into machine learning frameworks, building benchmarking pipelines, and leading optimization work across AI and HPC systems.
Location: 100% remote within the United States
Salary: $130,000–$180,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design, develop, and optimize CUDA kernels for AI, deep learning, and scientific computing workloads.
- Profile GPU workloads with NVIDIA Nsight Systems, Nsight Compute, CUDA Profiler, and related tools.
- Optimize GPU memory management, kernel execution, occupancy, multi-GPU scaling, and distributed computing performance.
- Design distributed training and inference architectures using NCCL, MPI, CUDA-aware communication libraries, and high-performance networking.
- Develop custom GPU operators and optimized kernels for PyTorch, JAX, Triton, TensorFlow, and similar frameworks.
- Build benchmarking and performance regression pipelines, collaborate with AI and software teams, and mentor engineers.
Requirements
- 10+ years of professional experience in GPU programming, high-performance computing, or parallel computing.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline.
- Expert proficiency in CUDA C/C++, GPU architecture, and massively parallel programming.
- Extensive experience with NCCL, MPI, CUDA-aware MPI, and distributed GPU communication frameworks.
- Experience integrating custom GPU kernels with PyTorch, TensorFlow, JAX, Triton, or similar machine learning frameworks.
- Strong C/C++ programming, debugging, profiling, analytical, collaboration, and technical leadership skills.
Nice to have
- Experience with CUTLASS, TensorRT, FasterTransformer, vLLM, DeepSpeed, or similar GPU optimization frameworks.
- Knowledge of LLVM, MLIR, compiler optimization, or code generation technologies.
- Experience with distributed AI training, model parallelism, pipeline parallelism, and inference optimization.
- Familiarity with AWS, Microsoft Azure, or Google Cloud Platform GPU infrastructure.
- Open-source contributions, research publications, patents, technical presentations, or experience with AMD ROCm and Intel accelerator technologies.
Culture & Benefits
- Full-time direct W2 employment.
- Career growth opportunities within an established consulting and software development organization.
- Work remotely within the United States.
- U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply.
- New H-1B visa petitions are not sponsored for this position.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Tenstorrent
10 дней назад
Staff Field Application Engineer (AI)
100 000 - 500 000$
10 дней назад
Founding GPU Engineer (AI)
11 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$
11 дней назад
Senior AI Applied Scientist II (Medtech)
150 000 - 177 000CAD
10 дней назад
Principal AI Engineer
175 000 - 200 000$
11 дней назад
Software Engineer (AI)
111 662 - 145 000$