Назад
Company hidden
1 день назад

Machine Learning Infrastructure Engineer (AI)

105 000 - 143 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Infrastructure Engineer (AI): Building and operating high-performance inference platforms for serving large machine learning models in production with an accent on request routing, batching, caching, autoscaling, GPU utilization, and observability. Focus on designing distributed serving systems, optimizing latency, throughput, cost, and quality, and supporting reliable AI APIs at scale.

Location: 100% remote within the United States

Salary: $105,000–$143,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design, build, and operate high-performance inference platforms for large machine learning models in production.
  • Develop request routing, batching, caching, autoscaling, and GPU utilization strategies for diverse model workloads.
  • Operate high-throughput, low-latency distributed services at scale.
  • Implement end-to-end observability using metrics, tracing, and structured logging.
  • Apply performance engineering and capacity planning to balance latency, throughput, cost, and quality.
  • Support reliable external-facing AI APIs and respond to production incidents.

Requirements

  • 6+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Hands-on experience with LLM or large-model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, accelerator utilization, Kubernetes, autoscaling, and modern cloud platforms.
  • Strong communication and incident response skills.

Nice to have

  • Open-source contributions to model-serving infrastructure.
  • Experience with multi-region or globally distributed AI serving.
  • Knowledge of model quantization, distillation, compression, or FinOps for AI workloads.
  • Experience supporting external-facing AI APIs at scale.

Culture & Benefits

  • Full-time direct W2 employment.
  • Fully remote work within the United States.
  • Career growth opportunities within an established organization.
  • U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates are encouraged to apply.
  • New H-1B visa petitions cannot be sponsored for this position.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →