обновлено 4 часа назад
MLOps Engineer (AI)
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
MLOps Engineer (AI serving): Building and operating high-performance inference platforms for LLMs, vision models, and recommendation systems with an accent on request routing, batching, GPU utilization, autoscaling, and observability. Focus on optimizing latency, throughput, cost, and quality through distributed systems engineering, deployment automation, multi-tenant controls, and reliable production operations.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate model-serving platforms for LLMs, vision models, and recommendation systems.
- Optimize inference performance through continuous batching, paged attention, speculative decoding, request multiplexing, caching, and prompt deduplication.
- Build multi-tenant routing, rate limiting, quality-of-service, autoscaling, and capacity-management systems.
- Tune GPU utilization, memory management, and KV-cache strategies for large-model serving workloads.
- Develop deployment workflows with canary releases, shadow testing, automated rollback, and end-to-end observability.
- Operate incident response for high-availability AI services and collaborate with ML and product teams on model releases, security controls, and reliability improvements.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- 6+ years of experience in distributed systems, infrastructure, or ML platform engineering.
- Strong Python proficiency and experience with a systems language such as Go, Rust, or C++.
- Experience operating high-throughput, low-latency production services and using LLM inference frameworks such as vLLM or TensorRT-LLM.
- Strong understanding of GPU architecture, memory hierarchies, accelerator utilization, performance engineering, and capacity planning.
- Familiarity with Kubernetes, autoscaling, cloud platforms, observability stacks, communication, and incident response.
Nice to have
- Open-source contributions to model-serving infrastructure.
- Experience with multi-region or globally distributed AI serving.
- Knowledge of model quantization, distillation, compression, or FinOps for AI workloads.
- Experience supporting external-facing AI APIs at scale.
Culture & Benefits
- Full-time direct W-2 employment.
- Career growth opportunities within an established technology consulting and software development organization.
- Work on cloud, AI, data, and enterprise solutions.
Hiring process
- Submit a resume for consideration.
- Applicants must be U.S. citizens, Green Card holders, EAD holders, or H-1B transfer candidates; new H-1B visa petitions cannot be sponsored.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Staff Software Engineer, Agentic AI
151 000 - 297 000$
2 дня назад
Advanced AI Engineer
103 000 - 155 000$
Baseten
3 дня назад
Software Engineer (AI)
180 000 - 360 000$
4 дня назад
Principal AI Engineer
1 день назад
Staff Software Engineer - Customer Facing Applied AI
158 000 - 248 000$
4 дня назад
Agentic AI Engineer
150 000 - 180 000$