Назад
Company hidden
4 дня назад

GPU Software Engineer (CUDA)

100 000 - 175 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
GPU Software Engineer (CUDA) (AI/HPC): Designing and optimizing high-performance CUDA kernels, GPU libraries, and distributed workloads for AI training, inference, scientific computing, and high-throughput data processing with an accent on memory hierarchy, mixed-precision computation, and multi-GPU performance. Focus on profiling production GPU systems, building custom operators and fused kernels, and contributing to compiler-level optimizations across ML frameworks and accelerator code generation.

Location: 100% remote within the United States

Salary: $100,000–$175,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design and implement high-performance CUDA kernels for AI and high-performance computing workloads.
  • Profile and optimize GPU code using Nsight Systems, Nsight Compute, and CUDA profiling tools.
  • Tune memory access, occupancy, register usage, shared memory, HBM, and L2 utilization.
  • Develop optimized libraries, custom operators, and fused kernels for PyTorch, JAX, or Triton.
  • Optimize multi-GPU and multi-node training with NCCL, RDMA, MPI, and high-performance networking.
  • Build benchmarks and regression tests, evaluate new GPU architectures, document tuning decisions, and mentor engineers.

Requirements

  • Must be based in the United States for this remote position.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of GPU programming and performance engineering experience.
  • Deep expertise in CUDA C/C++, GPU architectures, memory hierarchies, and execution models.
  • Production experience profiling and optimizing GPU workloads, including familiarity with NCCL, MPI, and high-performance interconnects.
  • Experience integrating custom kernels into ML frameworks, with strong systems programming, linear algebra, numerical methods, communication, and collaboration skills.

Nice to have

  • Experience with Triton, CUTLASS, TensorRT, FasterTransformer, or vLLM internals.
  • Exposure to LLVM or MLIR compiler infrastructure.
  • Open-source contributions to GPU or ML performance libraries.
  • Experience with large-scale distributed training infrastructure.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Remote work within the United States.
  • Collaboration with product, design, engineering, operations, business, research, and ML teams.
  • Opportunities for career growth, technical leadership, code review, design review, and mentorship.
  • New H-1B visa petitions are not sponsored; U.S. citizens, green card holders, EAD holders, and H-1B transfer candidates may apply.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →