14 часов назад
ML Performance Engineer (AI)
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Performance Engineer (AI): Optimizing end-to-end training and inference pipelines for large neural network systems with an accent on throughput, latency, cost efficiency, and GPU performance. Focus on distributed training, LLM serving, compiler-level optimization, custom kernels, and rigorous benchmarking across production workloads.
Location: 100% remote within the United States; the posting also lists Plymouth, MN.
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Profile and optimize AI training and inference pipelines for throughput, latency, and cost.
- Identify bottlenecks across data loading, model computation, communication, memory, storage, and distributed systems.
- Implement model compression, quantization, sparsity, pruning, KV-cache optimization, continuous batching, and speculative decoding.
- Optimize distributed training and inference using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
- Drive compiler and kernel-level improvements with Triton, XLA, TorchInductor, TVM, and related technologies.
- Build benchmark and regression frameworks, evaluate new hardware and software, and share performance playbooks across engineering teams.
Requirements
- Six or more years of experience in performance engineering, ML systems, or HPC.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
- Strong proficiency in Python and C++.
- Hands-on experience optimizing deep learning workloads on modern GPUs.
- Deep understanding of distributed training, inference, memory hierarchies, communication primitives, and parallelism strategies.
- Experience with CPU, GPU, and distributed-system profiling, along with strong measurement, debugging, analytical, communication, and collaboration skills.
Nice to have
- Production-scale LLM inference optimization experience.
- Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
- Custom kernel authoring with Triton or CUTLASS.
- FinOps experience for AI workloads.
- Publications or talks on AI systems performance.
Culture & Benefits
- Full-time direct W-2 employment.
- Collaboration with product, design, engineering, operations, and business stakeholders.
- Opportunities to contribute to production AI systems, engineering standards, code reviews, design reviews, and mentorship.
- Career growth within an established technology consulting and software development organization.
- Applicants must be authorized to work in the United States; new H-1B visa petitions cannot be sponsored.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Software Engineer (AI/ML Performance)
147 900 - 220 000$
7 дней назад
Staff ML Engineer (AI)
130 000 - 190 000$
Anthropic
4 дня назад
Performance Engineer (AI)
280 000 - 850 000$
3 дня назад
HPC and AI Performance Engineer (HPC/AI)
105 500 - 243 000$
Baseten
4 дня назад
Software Engineer (AI)
180 000 - 360 000$
7 дней назад
Engineering Leader (AI)
170 000 - 220 000$