Назад
Company hidden
обновлено 12 часов назад

AI Platform Engineer

130 000 - 180 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Platform Engineer (LLM Serving/Cloud Infrastructure): Building and operating enterprise-scale AI inference and model-serving platforms with an accent on distributed systems, Kubernetes, GPU optimization, and cloud-native infrastructure. Focus on optimizing inference latency and throughput, designing autoscaling and request routing, implementing MLOps lifecycle management, and ensuring observability, security, and high availability.

Location: 100% remote within the Continental United States

Salary: $130,000–$180,000 annually, based on experience.

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design, build, and maintain scalable AI inference and model-serving platforms for production environments.
  • Architect highly available cloud-native infrastructure for LLMs, foundation models, and machine learning services.
  • Optimize inference latency, throughput, GPU utilization, memory management, workload orchestration, and request routing.
  • Implement model deployment, versioning, rollback, lifecycle management, monitoring, logging, tracing, and alerting.
  • Build caching, API gateway, authentication, authorization, security, and high-availability solutions.
  • Collaborate with AI researchers, ML engineers, DevOps teams, and software engineers while mentoring engineers and driving infrastructure optimization.

Requirements

  • Must be based in the Continental United States.
  • 10+ years of professional experience in distributed systems, infrastructure, cloud platforms, or machine learning platform engineering.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Artificial Intelligence, or a related technical discipline.
  • Strong Python skills and proficiency in Go, Rust, or C++.
  • Experience with LLM serving, model inference optimization, vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar frameworks.
  • Expertise in Kubernetes, Docker, cloud platforms, CUDA, NVIDIA GPUs, distributed systems, networking, observability, and security.

Nice to have

  • Experience with multi-region AI platforms and globally distributed inference services.
  • Knowledge of quantization, pruning, compression, speculative decoding, KV cache optimization, and mixed-precision inference.
  • Experience with MLOps, GitOps, Terraform, Bicep, CloudFormation, and CI/CD automation.
  • Familiarity with Istio, Linkerd, API gateways, event-driven architectures, FinOps, and enterprise AI governance.
  • Open-source contributions, technical publications, patents, conference presentations, or experience supporting large-scale AI APIs.

Culture & Benefits

  • Full-time direct W2 employment.
  • Career growth within an established technology consulting and software development organization.
  • Work remotely within the Continental United States.
  • U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates are eligible to apply.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →