Назад
Company hidden
обновлено 5 дней назад

AI Optimization Engineer

100 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Optimization Engineer (GPU/Distributed AI Systems): Optimizing training and inference workloads for large neural networks with an accent on throughput, latency, cost, GPU architecture, and memory management. Focus on profiling CPU, GPU, and distributed systems, designing kernel and compiler-level optimizations, and improving production-scale AI performance.

Location: 100% remote within the United States

Salary: $100,000 annually

Company

Technology consulting and software development organization delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Optimize training and inference workloads for large neural network systems to increase throughput, reduce latency, and control costs.
  • Improve performance across the stack, from low-level kernels and compiler optimizations to distributed system configuration.
  • Analyze GPU architecture, model parallelism, memory management, communication primitives, and parallel execution strategies.
  • Profile and measure CPU, GPU, and distributed workloads to guide data-driven optimization decisions.
  • Collaborate with product, design, engineering, operations, and business stakeholders to turn ambiguous requirements into production-ready solutions.
  • Contribute through code reviews, design reviews, mentoring, and delivery of reliable production systems.

Requirements

  • Must be eligible to work in the United States as a U.S. citizen, Green Card holder, EAD holder, or H-1B transfer candidate.
  • New H-1B visa petitions cannot be sponsored for this position.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • 6+ years of experience in performance engineering, ML systems, or high-performance computing.
  • Strong proficiency in Python and C++, with hands-on experience optimizing deep learning workloads on modern GPUs.
  • Experience with distributed training and inference, profiling tools, model compression, debugging, measurement, and analytical reasoning.

Nice to have

  • Production-scale LLM inference optimization experience.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Experience authoring custom kernels with Triton or CUTLASS.
  • Experience with FinOps for AI workloads.
  • Publications or talks on AI systems performance.

Culture & Benefits

  • Fully remote work within the United States.
  • Full-time direct W-2 employment.
  • Career growth opportunities within an established organization.
  • Cross-functional collaboration with product, design, engineering, operations, and business stakeholders.
  • Equal employment opportunity and a workplace free from discrimination and harassment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →