Назад
Company hidden
обновлено 4 дня назад

AI Performance Engineer

75 000 - 100 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Performance Engineer (Python/C++/GPU): Optimizing training and inference workloads for large neural network systems with an accent on throughput, latency, cost, and distributed execution. Focus on low-level kernel optimization, GPU and memory performance, model parallelism, compiler-level optimization, and production-scale measurement.

Location: 100% remote within the United States

Salary: $75,000–$100,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Optimize training and inference workloads for large neural network systems to increase throughput, reduce latency, and control costs.
  • Improve performance across the stack, from low-level kernels and GPU execution to distributed system configuration.
  • Apply instrumentation, profiling, measurement, and debugging to make data-driven optimization decisions.
  • Design and tune model parallelism, memory management, communication primitives, and distributed training and inference systems.
  • Collaborate with product, design, engineering, operations, and business stakeholders to translate ambiguous requirements into production-ready solutions.
  • Contribute through code reviews, design reviews, technical communication, and mentorship of junior engineers.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Experience with profiling tools across CPU, GPU, and distributed systems, plus strong measurement, debugging, and analytical skills.
  • U.S. work authorization is required; new H-1B visa petitions cannot be sponsored.

Nice to have

  • Experience optimizing LLM inference at production scale.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Experience authoring custom kernels in Triton or CUTLASS.
  • Experience with FinOps for AI workloads.
  • Publications or talks on AI systems performance.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Opportunity to work on cloud, AI, data, and enterprise solutions.
  • Cross-functional collaboration with product, design, engineering, operations, and business stakeholders.
  • Career growth opportunities within an established organization.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →