Назад
2 дня назад

Software Engineer, Model Serving System (AI)

119 800 - 304 200$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer, Model Serving System (AI): Building and optimizing end-to-end LLM model serving systems for reliable production deployment on Microsoft Azure with an accent on distributed inference, GPU performance, benchmarking, and observability. Focus on integrating adapted models, improving latency and throughput, and solving complex deployment, caching, and evaluation challenges at scale.

Location: Mountain View, California, or Redmond, Washington, United States

Salary: USD $119,800–$234,700 per year for IC4 roles and USD $142,800–$274,800 per year for IC5 roles across the U.S.; location-specific ranges may reach USD $304,200 per year.

Company

Microsoft AI develops large language models and AI services delivered through the Microsoft Azure cloud platform.

What you will do

  • Integrate, optimize, and deploy LLMs in production model-serving environments.
  • Implement and evaluate serving features such as content provenance and watermarking while minimizing latency.
  • Profile GPU clusters and distributed serving systems against production SLOs including TTFT, TPOT, throughput, and QPS.
  • Support multi-LoRA and other adapted-model serving scenarios, including packaging and validation tooling.
  • Design realistic LLM benchmarks, improve endpoint evaluation workflows, and identify performance bottlenecks.
  • Build dashboards and telemetry for utilization, latency, cache efficiency, and SLA compliance.

Requirements

  • Bachelor's degree in Computer Science or a related technical field and 5+ years of technical engineering experience, or equivalent experience.
  • Strong coding experience in Python, C++, or similar languages.
  • Experience with LLM inference and serving systems such as vLLM, SGLang, TensorRT-LLM, or custom inference engines.
  • Strong systems and debugging skills, including performance analysis on GPU clusters.
  • Experience with distributed serving concepts including TP, PP, DP, KV cache, and speculative decoding.
  • Ability to work in the listed United States locations.

Nice to have

  • Master's degree and 5+ years of relevant experience.
  • Experience with Azure ML, Foundry, Kubernetes, or containerized model deployments.
  • Knowledge of multi-adapter serving, disaggregated prefill/decode, or large-scale KV cache management.
  • Experience with Grafana, Prometheus, Kusto, or equivalent observability tools.
  • Knowledge of GPU topology, memory management, modern accelerator hardware, and LLM benchmarking.

Culture & Benefits

  • High-ownership engineering work spanning platform, API, research, and engine teams.
  • Focus on making AI models shippable, measurable, and evaluable in production.
  • Benefits and additional compensation may be available depending on the role.
  • Applications are accepted on an ongoing basis until the position is filled.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →