4 дня назад
ML Performance Engineer (AI)
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Performance Engineer (AI): Optimizing end-to-end training and inference pipelines for large neural network systems with an accent on GPU performance, distributed computing, memory management, and compiler-level optimization. Focus on reducing latency and cost, implementing LLM serving techniques, building benchmark frameworks, and translating AI systems research into production improvements.
Location: 100% remote within the United States; the posting also lists Renton, Washington.
Salary: $100,000–$150,000 annually.
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
- Identify bottlenecks across data loading, model computation, communication, memory, and storage access.
- Implement quantization, sparsity, pruning, tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
- Optimize LLM serving with FlashAttention, paged attention, KV cache optimization, continuous batching, and speculative decoding.
- Drive compiler and kernel-level improvements using Triton, XLA, TorchInductor, or TVM.
- Build benchmark and regression frameworks, evaluate hardware and software, and share performance tuning practices across engineering teams.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
- Six or more years of experience in performance engineering, ML systems, or HPC.
- Strong proficiency in Python and C++.
- Hands-on experience optimizing deep learning workloads on modern GPUs.
- Deep understanding of distributed training, inference, profiling, memory hierarchies, communication primitives, and parallelism strategies.
- Applicants must be authorized to work in the United States; new H-1B visa petitions cannot be sponsored.
Nice to have
- Production-scale LLM inference optimization experience.
- Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
- Custom kernel authoring with Triton or CUTLASS.
- FinOps experience for AI workloads.
- Publications or talks on AI systems performance.
Culture & Benefits
- Fully remote work within the United States.
- Direct full-time W-2 employment.
- Collaboration with product, design, engineering, operations, and business stakeholders.
- Opportunities for code review, design review, mentorship, and career growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Machine Learning Performance Engineer (Inference)
200 000 - 300 000$
9 дней назад
HPC & AI Performance Engineer (HPC/AI)
62 900 - 145 300$
5 дней назад
Machine Learning Engineer III (AI)
141 400 - 190 700$
8 дней назад
Senior ML Systems Engineer (AI)
174 900 - 261 300$
VIA
5 дней назад
Senior Data Analytics Engineer (AI)
140 000 - 160 000$
Freedom Travel
9 дней назад
ML Engineer / Research Engineer
150 000 - 350 000₽