Назад
Company hidden
5 дней назад

ML Performance Engineer (AI)

100 000 - 150 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Performance Engineer (AI): Optimizing end-to-end training and inference pipelines for large neural network systems with an accent on throughput, latency, cost efficiency, and GPU performance. Focus on distributed training, model compression, LLM serving, compiler-level optimization, and building rigorous benchmark and regression frameworks.

Location: 100% remote within the United States; the posting also lists Bothell, WA.

Salary: $100,000–$150,000 annually.

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
  • Identify bottlenecks across data loading, model computation, communication, and memory.
  • Implement quantization, sparsity, pruning, KV cache optimization, continuous batching, and speculative decoding.
  • Optimize distributed training with tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Drive compiler and kernel optimizations using Triton, XLA, TorchInductor, or TVM.
  • Build benchmark suites and regression frameworks, evaluate new hardware and software, and document performance playbooks.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference, profiling, memory hierarchies, communication primitives, and parallelism.
  • U.S. work authorization is required; the role is open to U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates.

Nice to have

  • Production-scale LLM inference optimization experience.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Custom kernel authoring with Triton or CUTLASS.
  • FinOps experience for AI workloads.
  • Publications or talks on AI systems performance.

Culture & Benefits

  • Full-time direct W2 employment.
  • Collaboration with product, design, engineering, operations, and business stakeholders.
  • Opportunities to contribute to production AI systems and broader ML framework improvements.
  • Mentorship, code reviews, design reviews, and knowledge sharing across engineering teams.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →