Назад
Company hidden
обновлено 2 дня назад

ML Systems Engineer (AI)

145 000 - 165 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Systems Engineer (AI): Design, build, and operate high-performance inference platforms for serving large machine learning models in production with an accent on request routing, batching, caching, autoscaling, GPU utilization, and observability. Focus on shipping distributed serving systems at scale, optimizing latency, throughput, cost, and quality, and supporting reliable AI workloads.

Location: 100% remote within the United States

Salary: $145,000–$165,000 annually

Company

Cloud, AI, data, and enterprise technology consulting and software development services are delivered across the United States.

What you will do

  • Design, build, and operate high-performance inference platforms for large machine learning models in production.
  • Develop request routing, batching, caching, and autoscaling capabilities for diverse model workloads.
  • Optimize GPU utilization, latency, throughput, cost, and serving quality.
  • Implement end-to-end observability through metrics, tracing, and structured logging.
  • Support reliable, high-throughput, low-latency services and participate in incident response.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 6+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Experience operating high-throughput, low-latency production services and using LLM or large-model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, accelerator utilization, performance engineering, and capacity planning.
  • Familiarity with Kubernetes, autoscaling, modern cloud platforms, observability stacks, communication, and incident response.

Nice to have

  • Open-source contributions to model-serving infrastructure.
  • Experience with multi-region or globally distributed AI serving.
  • Familiarity with model quantization, distillation, and compression techniques.
  • Exposure to FinOps for AI workloads and cost-efficient serving design.
  • Experience supporting external-facing AI APIs at scale.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Career growth opportunities within an established technology consulting and software development organization.
  • Applicants must be eligible to work in the United States.
  • New H-1B visa petitions cannot be sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates are encouraged to apply.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →