4 дня назад
ML Platform Engineer (AI)
100 000 - 160 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Platform Engineer (AI): Designing, building, and operating high-performance inference platforms for large machine learning models in production with an accent on distributed systems, GPU utilization, request routing, and observability. Focus on optimizing latency, throughput, cost, and quality through batching, autoscaling, caching, deployment automation, and reliable incident response.
Location: 100% remote within the United States
Salary: $100,000–$160,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate model-serving platforms for LLMs, vision models, and recommendation systems.
- Optimize inference performance through continuous batching, paged attention, speculative decoding, request multiplexing, caching, and prompt deduplication.
- Build multi-tenant routing, rate limiting, quality-of-service, autoscaling, and capacity-management systems.
- Tune GPU utilization, memory management, and KV-cache strategies for large-model serving.
- Integrate serving platforms with API gateways, identity systems, and observability stacks.
- Develop canary releases, shadow testing, automated rollback, incident response, security controls, and operational documentation.
Requirements
- 10+ years of experience in distributed systems, infrastructure, or ML platform engineering.
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Strong proficiency in Python and a systems language such as Go, Rust, or C++.
- Experience operating high-throughput, low-latency production services and using LLM or large-model inference frameworks such as vLLM or TensorRT-LLM.
- Strong understanding of GPU architecture, memory hierarchies, accelerator utilization, performance engineering, and capacity planning.
- Familiarity with Kubernetes, autoscaling, cloud platforms, observability, communication, and incident response.
Nice to have
- Open-source contributions to model-serving infrastructure.
- Experience with multi-region or globally distributed AI serving.
- Familiarity with model quantization, distillation, compression, or FinOps for AI workloads.
- Experience supporting external-facing AI APIs at scale.
Culture & Benefits
- Full-time direct W-2 employment.
- Career growth opportunities within an established technology consulting and software development organization.
- Equal employment opportunity workplace.
- New H-1B visa petitions are not sponsored; U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates may apply.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →