3 дня назад
AI Performance Engineer
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Performance Engineer (AI/LLM Systems): Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, cost, and memory efficiency with an accent on distributed training, model compression, GPU workloads, and compiler-level optimization. Focus on implementing advanced LLM serving techniques, building benchmark and regression frameworks, and translating AI systems research into measurable production performance gains.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, cost, and memory efficiency.
- Identify bottlenecks across data loading, model computation, communication, storage, and memory systems.
- Implement model compression, quantization, sparsity, pruning, KV cache optimization, continuous batching, and speculative decoding.
- Optimize distributed training and inference using tensor parallelism, pipeline parallelism, FSDP, ZeRO-style sharding, and advanced attention implementations.
- Drive compiler and kernel-level improvements with Triton, XLA, TorchInductor, TVM, and related technologies.
- Build benchmark and regression frameworks, evaluate hardware and software options, improve AI cost efficiency, and share performance practices across engineering teams.
Requirements
- Must be authorized to work in the United States as a U.S. citizen, Green Card holder, EAD holder, or H-1B transfer candidate.
- New H-1B visa petitions cannot be sponsored for this position.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
- 10+ years of experience in performance engineering, ML systems, or HPC.
- Strong proficiency in Python and C++, with hands-on experience optimizing deep learning workloads on modern GPUs.
- Deep understanding of distributed training and inference, profiling tools, memory hierarchies, communication primitives, parallelism, measurement, debugging, and analytical reasoning.
Nice to have
- Production-scale LLM inference optimization experience.
- Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
- Custom kernel authoring experience with Triton or CUTLASS.
- FinOps experience for AI workloads.
- Publications or talks on AI systems performance.
Culture & Benefits
- Full-time direct W-2 employment.
- Fully remote work within the United States.
- Career growth opportunities within an established technology consulting and software development organization.
- Collaboration with ML, platform, and broader engineering teams.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
AI Software Development Engineer (Neuromorphic Computing)
170 500 - 240 710$
4 дня назад
Engineering Leader (AI)
170 000 - 220 000$
4 дня назад
Research Engineer (AI)
190 000 - 240 000$
4 дня назад
Engineering Leader (AI)
90 000 - 120 000€
4 дня назад
Staff ML Engineer (AI)
130 000 - 190 000$
Databricks
10 дней назад
AI Engineer – Forward Deployed Engineering (U.S. Public Sector)
182 000 - 250 208$