Назад
Company hidden
6 дней назад

ML Platform Engineer (AI)

100 000 - 160 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Platform Engineer (AI): Building and operating high-performance inference platforms for serving large machine learning models in production with an accent on distributed systems, GPU utilization, request routing, and observability. Focus on optimizing latency, throughput, cost, and reliability through batching, autoscaling, caching, deployment automation, and incident response.

Location: 100% remote within the United States; the posting also lists Nashua, NH.

Salary: $100,000–$160,000 annually.

Company

Technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design and operate model-serving platforms for LLM, vision, and recommendation workloads.
  • Optimize inference with continuous batching, paged attention, speculative decoding, request multiplexing, caching, and prompt deduplication.
  • Build multi-tenant routing, rate limiting, quality-of-service, autoscaling, and capacity-management systems.
  • Tune GPU utilization, memory management, and KV-cache strategies for large-model serving.
  • Integrate serving platforms with API gateways, identity systems, observability tools, and security controls.
  • Develop canary releases, shadow testing, automated rollback, incident response, and reliability improvements.

Requirements

  • Must be authorized to work in the United States as a U.S. citizen, Green Card holder, EAD holder, or H-1B transfer candidate; new H-1B sponsorship is unavailable.
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 10+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Experience operating high-throughput, low-latency production services and using LLM or large-model inference frameworks such as vLLM or TensorRT-LLM.
  • Experience with GPU architecture, Kubernetes, cloud platforms, autoscaling, observability, performance engineering, capacity planning, communication, and incident response.

Nice to have

  • Open-source contributions to model-serving infrastructure.
  • Experience with multi-region or globally distributed AI serving.
  • Knowledge of model quantization, distillation, compression, FinOps, or external-facing AI APIs at scale.

Culture & Benefits

  • Full-time direct W-2 employment.
  • 100% remote work within the United States.
  • Career growth opportunities within an established organization.
  • Equal employment opportunity and anti-harassment commitments.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →