Назад
Company hidden
4 дня назад

ML Platform Engineer (AI)

100 000 - 160 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Platform Engineer (AI): Designing and operating high-performance inference platforms for serving LLMs, vision models, and recommendation systems in production with an accent on request routing, batching, autoscaling, GPU utilization, and observability. Focus on optimizing latency, throughput, cost, and quality through distributed systems engineering, multi-tenant serving, deployment automation, and high-availability incident response.

Location: 100% remote within the United States

Salary: $100,000–$160,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design and operate model-serving platforms for LLMs, vision models, and recommendation systems.
  • Optimize inference performance through continuous batching, paged attention, speculative decoding, request multiplexing, caching, and prompt deduplication.
  • Build multi-tenant routing, rate limiting, quality-of-service policies, autoscaling, and capacity-management systems.
  • Tune GPU utilization, memory management, and KV-cache strategies while balancing latency, throughput, cost, and quality.
  • Integrate serving platforms with API gateways, identity systems, observability tools, and security controls.
  • Develop canary releases, shadow testing, automated rollback, incident response procedures, and reliability improvements for high-availability AI services.

Requirements

  • Must be based in the United States.
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 10+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Experience operating high-throughput, low-latency production services and using LLM or large-model inference frameworks such as vLLM or TensorRT-LLM.
  • Experience with GPU architecture, Kubernetes, autoscaling, cloud platforms, observability stacks, performance engineering, capacity planning, and incident response.

Nice to have

  • Open-source contributions to model-serving infrastructure.
  • Experience with multi-region or globally distributed AI serving.
  • Familiarity with model quantization, distillation, compression, and FinOps for AI workloads.
  • Experience supporting external-facing AI APIs at scale.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Opportunity to work on production AI infrastructure and large-scale model serving.
  • Collaboration with ML and product teams on model releases and capability rollouts.
  • Career growth within an established technology consulting and software development organization.

Hiring process

  • Submit a resume for consideration.
  • U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates may apply; new H-1B visa petitions are not sponsored.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →