Назад
Company hidden
3 дня назад

Senior Software Engineer, Inference (AI)

144 000 - 273 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer, Inference (AI): Building and evolving the LLM model runtime and Kubernetes orchestration layer for enterprise AI inference on customer-owned infrastructure with an accent on engine integration, continuous batching, KV cache management, and distributed execution. Focus on optimizing tail latency and GPU utilization, implementing parallel and disaggregated serving, and solving complex multi-tier inference performance challenges.

Location: Hybrid, with an average of 2 days per week from an HPE office; other HPE site locations in the United States may be considered. Remote work options will be considered.

Salary: USD 144,000–273,000 annually in Colorado; USD 137,000–315,000 annually in North Carolina and Texas. Variable incentives may also be offered.

Company

hirify.global is a global edge-to-cloud technology company developing infrastructure and software for connecting, protecting, analyzing, and operating data and applications.

What you will do

  • Design, implement, and own LLM serving runtime components, including engine integration, continuous batching, KV cache management and reuse, and quantized execution.
  • Improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency.
  • Build distributed execution capabilities such as disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload.
  • Evaluate inference runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and recommend adoption where appropriate.
  • Develop Kubernetes orchestration features for model admission, GPU scheduling and partitioning, cache-aware routing, and autoscaling.
  • Resolve customer issues, conduct code and design reviews, mentor team members, and improve engineering practices.

Requirements

  • Hybrid work from an HPE office in the United States is required on average 2 days per week; remote options may be considered.
  • At least 8 years of software engineering experience, including 1–2 or more years working directly on LLM inference runtimes or production model serving.
  • Experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modifying engine internals.
  • Strong knowledge of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, speculative decoding, tensor and pipeline parallelism, and NCCL.
  • Advanced Kubernetes platform architecture knowledge, including operators, custom resources, controllers, and scheduling.
  • Strong Go and Python skills, with the ability to debug and profile C++/CUDA using tools such as Nsight; a Computer Science or related degree is required.

Nice to have

  • Upstream contributions to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe.
  • Experience with disaggregated prefill/decode serving, KV cache offload, RDMA, GPUDirect Storage, InfiniBand, or RoCE.
  • Experience with MIG, fractional GPU allocation, multi-tenant GPU isolation, or on-premises, air-gapped, and regulated enterprise software.

Culture & Benefits

  • Health, financial, and emotional wellbeing benefits for employees and their families.
  • Personal and professional development programs supporting career growth and technical expertise.
  • Inclusive workplace committed to accessibility and equal employment opportunities.
  • Flexibility to manage work and personal needs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →