Назад
Company hidden
3 дня назад

AI Performance Engineer

100 000 - 150 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Performance Engineer (AI/LLM Systems): Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, cost, and memory efficiency with an accent on distributed training, model compression, GPU workloads, and compiler-level optimization. Focus on implementing advanced LLM serving techniques, building benchmark and regression frameworks, and translating AI systems research into measurable production performance gains.

Location: 100% remote within the United States

Salary: $100,000–$150,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, cost, and memory efficiency.
  • Identify bottlenecks across data loading, model computation, communication, storage, and memory systems.
  • Implement model compression, quantization, sparsity, pruning, KV cache optimization, continuous batching, and speculative decoding.
  • Optimize distributed training and inference using tensor parallelism, pipeline parallelism, FSDP, ZeRO-style sharding, and advanced attention implementations.
  • Drive compiler and kernel-level improvements with Triton, XLA, TorchInductor, TVM, and related technologies.
  • Build benchmark and regression frameworks, evaluate hardware and software options, improve AI cost efficiency, and share performance practices across engineering teams.

Requirements

  • Must be authorized to work in the United States as a U.S. citizen, Green Card holder, EAD holder, or H-1B transfer candidate.
  • New H-1B visa petitions cannot be sponsored for this position.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • 10+ years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++, with hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference, profiling tools, memory hierarchies, communication primitives, parallelism, measurement, debugging, and analytical reasoning.

Nice to have

  • Production-scale LLM inference optimization experience.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Custom kernel authoring experience with Triton or CUTLASS.
  • FinOps experience for AI workloads.
  • Publications or talks on AI systems performance.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Fully remote work within the United States.
  • Career growth opportunities within an established technology consulting and software development organization.
  • Collaboration with ML, platform, and broader engineering teams.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →