3 дня назад
ML Platform Engineer (AI)
100 000 - 160 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Platform Engineer (AI): Designing and operating high-performance inference platforms for serving large machine learning models in production with an accent on distributed systems, GPU utilization, request routing, and observability. Focus on optimizing latency, throughput, cost, and quality through batching, autoscaling, caching, multi-tenant serving, and reliable deployment workflows.
Location: 100% remote within the United States
Salary: $100,000–$160,000 annually
Company
Technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate model-serving platforms for LLMs, vision models, and recommendation systems.
- Optimize inference performance through continuous batching, paged attention, speculative decoding, request multiplexing, caching, and prompt deduplication.
- Build multi-tenant routing, rate limiting, quality-of-service policies, autoscaling, and capacity management systems.
- Tune GPU utilization, memory management, and KV-cache strategies for large-model serving workloads.
- Integrate serving platforms with API gateways, identity systems, observability tools, and security controls.
- Develop canary releases, shadow testing, automated rollback, incident response, and reliability improvements for high-availability AI services.
Requirements
- 10+ years of experience in distributed systems, infrastructure, or ML platform engineering.
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Strong proficiency in Python and a systems language such as Go, Rust, or C++.
- Experience operating high-throughput, low-latency production services and working with LLM or large-model inference frameworks such as vLLM or TensorRT-LLM.
- Understanding of GPU architecture, memory hierarchies, accelerator utilization, Kubernetes, autoscaling, cloud platforms, and performance engineering.
- Experience with metrics, tracing, structured logging, capacity planning, communication, and incident response.
Nice to have
- Open-source contributions to model-serving infrastructure.
- Experience with multi-region or globally distributed AI serving and external-facing AI APIs at scale.
- Knowledge of model quantization, distillation, compression, and FinOps for AI workloads.
Culture & Benefits
- Full-time direct W-2 employment.
- Opportunity for career growth within an established technology organization.
- Work remotely within the United States.
- New H-1B visa petitions are not sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates may apply.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Waymo
4 дня назад
ML Engineer, Foundation Model Infrastructure
175 000 - 215 000$
5 дней назад
ML Platform Engineer (AI)
5 дней назад
ML Platform Engineer (AI)
4 дня назад
Staff AI Platform Engineer (Inference & Agentic Systems)
9 дней назад
ML Data & Platform Engineer (Speech AI)
5 дней назад