Назад
Company hidden
обновлено 2 дня назад

Model Serving Engineer (AI)

74 000 - 98 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Model Serving Engineer (AI): Design, build, and operate high-performance inference platforms for serving large machine learning models in production with an accent on request routing, batching, caching, autoscaling, GPU utilization, and observability. Focus on building distributed serving systems, optimizing latency, throughput, cost, and quality, and supporting reliable model workloads at scale.

Location: 100% remote within the United States

Salary: $74,000–$98,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design, build, and operate high-performance inference platforms for large machine learning models in production.
  • Develop serving capabilities for request routing, batching, caching, autoscaling, and GPU utilization.
  • Build end-to-end observability across diverse model workloads using metrics, tracing, and structured logging.
  • Optimize trade-offs between latency, throughput, cost, and quality in ML serving systems.
  • Apply distributed systems, performance engineering, and capacity planning practices to high-throughput services.
  • Support reliable external-facing AI APIs and respond to production incidents.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • At least 6 years of experience in distributed systems, infrastructure, or ML platform engineering; the position lists 7+ years of experience.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Experience operating high-throughput, low-latency production services.
  • Hands-on experience with LLM or large-model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, accelerator utilization, Kubernetes, autoscaling, cloud platforms, observability, and incident response.

Nice to have

  • Open-source contributions to model-serving infrastructure.
  • Experience with multi-region or globally distributed AI serving.
  • Knowledge of model quantization, distillation, compression, and FinOps for AI workloads.
  • Experience supporting external-facing AI APIs at scale.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Career growth opportunities within an established technology consulting and software development organization.
  • Equal employment opportunity and a workplace free from harassment and discrimination.
  • New H-1B visa petitions are not sponsored.
  • U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates are encouraged to apply.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →