Senior Software Engineer, Inference (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Hybrid, with an average of 2 days per week from an HPE office; other HPE site locations in the United States may be considered. Remote work options will be considered.
Salary: USD 144,000–273,000 annually in Colorado; USD 137,000–315,000 annually in North Carolina and Texas. Variable incentives may also be offered.
Company
is a global edge-to-cloud technology company developing infrastructure and software for connecting, protecting, analyzing, and operating data and applications.
What you will do
- Design, implement, and own LLM serving runtime components, including engine integration, continuous batching, KV cache management and reuse, and quantized execution.
- Improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency.
- Build distributed execution capabilities such as disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload.
- Evaluate inference runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and recommend adoption where appropriate.
- Develop Kubernetes orchestration features for model admission, GPU scheduling and partitioning, cache-aware routing, and autoscaling.
- Resolve customer issues, conduct code and design reviews, mentor team members, and improve engineering practices.
Requirements
- Hybrid work from an HPE office in the United States is required on average 2 days per week; remote options may be considered.
- At least 8 years of software engineering experience, including 1–2 or more years working directly on LLM inference runtimes or production model serving.
- Experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modifying engine internals.
- Strong knowledge of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, speculative decoding, tensor and pipeline parallelism, and NCCL.
- Advanced Kubernetes platform architecture knowledge, including operators, custom resources, controllers, and scheduling.
- Strong Go and Python skills, with the ability to debug and profile C++/CUDA using tools such as Nsight; a Computer Science or related degree is required.
Nice to have
- Upstream contributions to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe.
- Experience with disaggregated prefill/decode serving, KV cache offload, RDMA, GPUDirect Storage, InfiniBand, or RoCE.
- Experience with MIG, fractional GPU allocation, multi-tenant GPU isolation, or on-premises, air-gapped, and regulated enterprise software.
Culture & Benefits
- Health, financial, and emotional wellbeing benefits for employees and their families.
- Personal and professional development programs supporting career growth and technical expertise.
- Inclusive workplace committed to accessibility and equal employment opportunities.
- Flexibility to manage work and personal needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →