Назад
Company hidden
обновлено 4 дня назад

Model Optimization Engineer (AI)

150 000 - 175 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Model Optimization Engineer (AI): Optimizing training and inference workloads for large neural network systems with an accent on throughput, latency, cost, GPU architecture, and distributed systems. Focus on low-level kernel and compiler optimization, memory management, profiling, model parallelism, and shipping production-grade performance improvements.

Location: 100% remote within the United States

Salary: $150,000–$175,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Optimize throughput, latency, and cost across model training and inference workloads.
  • Improve large neural network systems across the stack, from low-level kernels to distributed systems.
  • Apply GPU architecture, model parallelism, memory management, and compiler-level optimization techniques.
  • Use instrumentation, profiling, measurement, and debugging to make data-driven optimization decisions.
  • Collaborate with product, design, engineering, operations, and business stakeholders to turn ambiguous requirements into production solutions.
  • Contribute through code reviews, design reviews, mentoring, and delivery of reliable production work.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference, profiling tools, model compression, memory hierarchies, communication primitives, and parallelism strategies.
  • Must be authorized to work in the United States as a U.S. citizen, Green Card holder, EAD holder, or H-1B transfer candidate; new H-1B sponsorship is not available.

Nice to have

  • Experience optimizing LLM inference at production scale.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Experience authoring custom kernels with Triton or CUTLASS.
  • Experience with FinOps for AI workloads.
  • Publications or talks on AI systems performance.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Career growth opportunities within an established organization.
  • Cross-functional collaboration with product, design, engineering, operations, and business stakeholders.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →