Назад
12 дней назад

Senior Evaluation Algorithm Engineer (AI)

Формат работы
remote (только Hong_kong)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
China
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Evaluation Algorithm Engineer (AI): Building end-to-end LLM evaluation systems for dialogue, financial trading, and other business scenarios with an accent on evaluation metrics, datasets, rubrics, and reliable model assessment. Focus on analyzing model failure modes, automating evaluation workflows, and translating findings into concrete improvements with algorithm, product, and data teams.

Location: Hong Kong / Remote

Company

Binance operates a global blockchain ecosystem spanning cryptocurrency trading, financial services, payments, education, research, and Web3 products.

What you will do

  • Design end-to-end LLM evaluation plans for dialogue, financial trading, and other business scenarios.
  • Build evaluation metrics, rubrics, datasets, benchmarks, annotation guidelines, and quality-control processes.
  • Analyze model capabilities and failure modes, then produce actionable recommendations for model improvement.
  • Automate and scale evaluation workflows by developing sustainable platforms and toolchains.
  • Collaborate with algorithm, product, and data teams to convert business and model objectives into evaluation standards and R&D directions.

Requirements

  • Master’s degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or a related field.
  • Strong algorithmic foundation and understanding of LLM principles, training, and fine-tuning.
  • Hands-on LLM evaluation experience in a large technology company, including evaluation for commercial deployment.
  • Knowledge of human evaluation, model-based automatic evaluation, LLM-as-a-judge methods, and metric computation.
  • Proficiency in Python with experience in evaluation automation, benchmark construction, platform development, data processing, scripting, and result analysis.
  • Strong business understanding and communication skills for driving cross-functional collaboration.

Nice to have

  • Experience evaluating dialogue systems, AI agents, or financial and trading LLMs.
  • Experience building AI training or evaluation data and annotation systems.
  • Familiarity with RLHF, reward models, or preference data.

Culture & Benefits

  • Work-from-home arrangement, with the arrangement potentially varying by business team and work nature.
  • Global organization with a flat structure and collaboration with international talent.
  • Fast-paced projects with autonomy and opportunities for continuous learning and career growth.
  • Competitive salary and company benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →