4 дня назад
AI Performance Engineer
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Performance Engineer (Python/C++/GPUs): Profiling and optimizing end-to-end AI training and inference pipelines with an accent on distributed workloads, model compression, memory efficiency, and compiler-level acceleration. Focus on improving LLM serving through KV cache optimization, continuous batching, speculative decoding, custom kernels, and rigorous benchmarking for measurable throughput, latency, and cost gains.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
- Identify bottlenecks across data loading, model computation, communication, memory, storage access, and distributed workloads.
- Implement model compression and inference optimizations including quantization, sparsity, pruning, FlashAttention, paged attention, KV cache optimization, continuous batching, and speculative decoding.
- Optimize distributed training with tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
- Drive compiler and kernel-level improvements using Triton, XLA, TorchInductor, TVM, or similar technologies.
- Build benchmark and regression frameworks, evaluate new hardware and software, document tuning practices, and collaborate with ML and platform engineering teams.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
- Six or more years of experience in performance engineering, ML systems, or HPC.
- Strong proficiency in Python and C++.
- Hands-on experience optimizing deep learning workloads on modern GPUs.
- Deep understanding of distributed training, inference, memory hierarchies, communication primitives, and parallelism strategies.
- Must be eligible to work in the United States; new H-1B visa petitions cannot be sponsored.
Nice to have
- Production-scale LLM inference optimization experience.
- Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
- Custom kernel authoring with Triton or CUTLASS.
- FinOps experience for AI workloads.
- Publications or talks on AI systems performance.
Culture & Benefits
- Full-time direct W-2 employment.
- Career growth opportunities within an established organization.
- Collaboration with ML, platform engineering, and broader engineering teams.
- Opportunity to translate AI systems research into production improvements.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
AI Infrastructure Engineer
170 500 - 315 490$
3 дня назад
Principal AI Engineer
180 000 - 200 000$
5 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$
3 дня назад
Principal GenAI Engineer (AI)
150 000 - 200 000$
5 дней назад
Principal AI Engineer
180 000 - 200 000$
2 дня назад
Applied AI Engineer
125 000 - 155 000$