Назад
Company hidden
обновлено 10 дней назад

Parallel Computing Engineer (CUDA)

130 000 - 180 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Parallel Computing Engineer (CUDA/HPC): Designing and optimizing GPU kernels, distributed training and inference architectures, and large-scale AI and scientific computing workloads with an accent on CUDA performance, GPU memory management, and multi-GPU scaling. Focus on profiling with NVIDIA Nsight tools, integrating custom operators into machine learning frameworks, building benchmarking pipelines, and leading optimization work across AI and HPC systems.

Location: 100% remote within the United States

Salary: $130,000–$180,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design, develop, and optimize CUDA kernels for AI, deep learning, and scientific computing workloads.
  • Profile GPU workloads with NVIDIA Nsight Systems, Nsight Compute, CUDA Profiler, and related tools.
  • Optimize GPU memory management, kernel execution, occupancy, multi-GPU scaling, and distributed computing performance.
  • Design distributed training and inference architectures using NCCL, MPI, CUDA-aware communication libraries, and high-performance networking.
  • Develop custom GPU operators and optimized kernels for PyTorch, JAX, Triton, TensorFlow, and similar frameworks.
  • Build benchmarking and performance regression pipelines, collaborate with AI and software teams, and mentor engineers.

Requirements

  • 10+ years of professional experience in GPU programming, high-performance computing, or parallel computing.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline.
  • Expert proficiency in CUDA C/C++, GPU architecture, and massively parallel programming.
  • Extensive experience with NCCL, MPI, CUDA-aware MPI, and distributed GPU communication frameworks.
  • Experience integrating custom GPU kernels with PyTorch, TensorFlow, JAX, Triton, or similar machine learning frameworks.
  • Strong C/C++ programming, debugging, profiling, analytical, collaboration, and technical leadership skills.

Nice to have

  • Experience with CUTLASS, TensorRT, FasterTransformer, vLLM, DeepSpeed, or similar GPU optimization frameworks.
  • Knowledge of LLVM, MLIR, compiler optimization, or code generation technologies.
  • Experience with distributed AI training, model parallelism, pipeline parallelism, and inference optimization.
  • Familiarity with AWS, Microsoft Azure, or Google Cloud Platform GPU infrastructure.
  • Open-source contributions, research publications, patents, technical presentations, or experience with AMD ROCm and Intel accelerator technologies.

Culture & Benefits

  • Full-time direct W2 employment.
  • Career growth opportunities within an established consulting and software development organization.
  • Work remotely within the United States.
  • U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply.
  • New H-1B visa petitions are not sponsored for this position.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →