Назад
Company hidden
обновлено 4 дня назад

High-Performance Computing Engineer (CUDA)

130 000 - 150 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
High-Performance Computing Engineer (CUDA): Design and optimize compute-intensive GPU workloads for AI training, inference, scientific computing, and high-throughput data processing with an accent on CUDA C/C++, GPU architecture, distributed computing, and production performance engineering. Focus on profiling GPU workloads, integrating custom kernels into ML frameworks, optimizing memory and execution behavior, and mentoring engineers through code and design reviews.

Location: 100% remote within the United States

Salary: $130,000–$150,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design and optimize compute-intensive workloads for modern GPU and accelerator platforms.
  • Improve performance for AI training, inference, scientific computing, and high-throughput data processing.
  • Profile production GPU workloads and deliver measurable performance improvements.
  • Integrate custom kernels into machine learning frameworks and work with distributed computing technologies.
  • Collaborate with product, design, engineering, operations, and business stakeholders to turn ambiguous requirements into production-ready solutions.
  • Contribute through code reviews, design reviews, technical mentorship, and strong engineering practices.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • 6+ years of experience in GPU programming and performance engineering.
  • Deep expertise in CUDA C/C++, GPU programming models, modern GPU architectures, memory hierarchies, and execution models.
  • Hands-on experience profiling and optimizing GPU workloads in production.
  • Familiarity with NCCL, MPI, high-performance interconnects, custom ML kernels, modern C++ systems programming, linear algebra, and numerical methods.
  • Applicants must be U.S. citizens, Green Card holders, EAD holders, or eligible for an H-1B transfer; new H-1B petitions cannot be sponsored.

Nice to have

  • Experience with Triton, CUTLASS, TensorRT, FasterTransformer, or vLLM internals.
  • Exposure to LLVM or MLIR compiler infrastructure.
  • Open-source contributions to GPU or ML performance libraries.
  • Experience with large-scale distributed training infrastructure.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Established organization with opportunities for career growth.
  • Collaborative work with research and engineering teams.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →