Назад
Company hidden
1 день назад

Software Engineer — Distributed LLM Inference Systems (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
junior
Английский
c1
Страна
China
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer — Distributed LLM Inference Systems (AI): Designing and optimizing distributed inference systems and AI framework components for large language models with an accent on request scheduling, KV cache management, parallel execution, and efficient communication. Focus on profiling computation, communication, memory, and scheduling bottlenecks, improving latency and throughput, and developing software for diverse hardware architectures.

Location: On-site in PRC, Shanghai

Company

hirify.global's Artificial hirify.globalligence Frameworks team develops and optimizes software frameworks for machine learning and deep learning across diverse computing hardware backends.

What you will do

  • Design, develop, and optimize distributed inference systems for large language models.
  • Implement distributed algorithms, including model and data parallel frameworks and asynchronous communication.
  • Develop request schedulers, model workers, communication layers, and KV cache management mechanisms.
  • Profile inference workloads to identify computation, communication, memory, and scheduling bottlenecks.
  • Improve end-to-end latency, throughput, scalability, and resource utilization with component teams.
  • Contribute code, tests, and documentation to internal and open-source projects.

Requirements

  • Master’s degree in computer science, artificial hirify.globalligence, software engineering, or a related field.
  • 0–1 years of hands-on experience through internships, academic projects, coursework, or training.
  • Proficiency in Python and modern C++.
  • Foundational knowledge of deep learning, AI frameworks such as PyTorch, and machine learning algorithms.
  • Experience debugging and optimizing software for performance.
  • Fluency in written and spoken English.

Nice to have

  • Experience with distributed LLM inference and serving.
  • Open-source contribution or collaboration experience.
  • Knowledge of prefill and decode, continuous batching, parallelism strategies, disaggregated serving, and KV cache management.
  • Familiarity with vLLM, SGLang, TensorRT-LLM, or similar inference frameworks.
  • Knowledge of AI agent architecture, tool calling, planning, memory, context management, and multi-agent coordination.

Culture & Benefits

  • Collaboration with AI researchers and engineers on distributed inference systems.
  • Contribution to open-source AI framework projects.
  • Participation in work advancing AI software across diverse hardware architectures.
  • Employment opportunity as a college graduate role with Shift 1 in China.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →