Назад
Company hidden
11 часов назад

Machine Learning Engineer (AI)

Формат работы
remote (только China)
Тип работы
fulltime
Английский
b2
Страна
China
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer (AI) (LLM systems): Building production-grade ML pipelines, inference systems, evaluation tooling, and deployment infrastructure for a proactive smart assistant with an accent on long-running workflows, persistent context, and reliable real-world task completion. Focus on fine-tuning transformer models, optimizing GPU-based inference, and balancing latency, cost, reliability, safety, and model behavior.

Location: Remote from China; address in Beijing, China. Interviews may be conducted virtually or onsite.

Company

hirify.global is building a proactive AI smart assistant for conversations, errands, organization, and everyday workflows.

What you will do

  • Build and own end-to-end ML pipelines across data, training, evaluation, inference, and deployment.
  • Fine-tune and adapt transformer-based models using LoRA, QLoRA, SFT, DPO, and distillation.
  • Architect scalable inference systems while balancing latency, cost, reliability, and safety.
  • Design data systems for synthetic and real-world training data and implement evaluation for performance, robustness, safety, and bias.
  • Own production deployment, including GPU optimization, memory efficiency, quantization, latency reduction, and scaling.
  • Collaborate with application engineering to integrate ML systems into backend, mobile, and desktop products.

Requirements

  • Strong background in deep learning and transformer-based architectures.
  • Experience training, fine-tuning, or deploying large-scale ML models in production.
  • Proficiency with a modern ML framework such as PyTorch or JAX.
  • Experience with distributed training and inference frameworks such as DeepSpeed, FSDP, Megatron, ZeRO, or Ray.
  • Strong software engineering fundamentals and experience building robust, maintainable, production-grade systems.
  • Experience with GPU optimization, including memory efficiency, quantization, and mixed precision, plus the ability to own ambiguous ML systems end-to-end.

Nice to have

  • Experience with vLLM, TensorRT-LLM, or FasterTransformer.
  • Open-source contributions to ML or systems libraries.
  • Background in scientific computing, compilers, or GPU kernels.
  • Experience with RLHF pipelines, multimodal or diffusion models, or large-scale data processing with Apache Arrow, Spark, or Ray.

Culture & Benefits

  • Work with a small, high-talent-density, hands-on team.
  • Make decisions collectively and ship improvements quickly through iterative learning.
  • Operate with autonomy, structure, and pragmatic judgment under real production constraints.
  • Prompt hiring decisions after the interview process.

Hiring process

  • Complete 3, and no more than 4, interviews if selected for further consideration.
  • Applications are evaluated by technical team members.
  • Interviews take place via virtual meetings and/or onsite.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →