Назад
Company hidden
2 месяца назад

Cloud Orchestration Engineer (AI)

200 000 - 400 000SGD
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Cloud Orchestration Engineer (AI): Building the operational backbone for reliable vLLM deployments at massive scale with an accent on cluster management, deployment automation, and production monitoring. Focus on designing custom Kubernetes operators, managing large GPU clusters, and improving observability, debuggability, and recoverability across cloud and on-premise infrastructure.

Location: Singapore; on-site

Salary: S$200,000–S$400,000 annually plus equity

Company

hirify.global was founded by the creators and core maintainers of vLLM to make AI inference cheaper and faster and help grow vLLM as an AI inference engine.

What you will do

  • Build the operational backbone for reliable vLLM deployments at massive scale.
  • Design systems for cluster management and deployment automation.
  • Develop production monitoring that makes deployments observable, debuggable, and recoverable.
  • Enable teams to serve AI models efficiently across cloud and on-premise infrastructure.
  • Improve operational reliability and resource management for ML inference systems.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a related field.
  • Strong experience with Kubernetes and container orchestration at scale.
  • Experience designing and implementing custom Kubernetes operators.
  • Proficiency in Python, Rust, or Go and infrastructure-as-code tools such as Terraform and Helm.
  • Experience managing GPU clusters and debugging hardware issues.
  • Ability to work across AWS, GCP, Azure, and on-premise infrastructure.

Nice to have

  • Experience with Ray or Slurm.
  • Knowledge of GPU scheduling, multi-tenancy, and resource optimization.
  • Familiarity with vLLM deployment patterns and configuration.
  • Experience improving operational reliability for ML systems.
  • Experience deploying inference systems on GPU clusters with 1,000 or more GPUs.

Culture & Benefits

  • Full-time position in the Research & Engineering department.
  • Medical, dental, and vision coverage.
  • Visa sponsorship is available on a case-by-case basis.
  • Equity included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →