Назад
Company hidden
4 дня назад

Applied Scientist, Agent Evaluation & Adaptive Model Routing (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
Singapore/US/Norway +4 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Applied Scientist, Agent Evaluation & Adaptive Model Routing (AI): Building evaluation and decision systems for agentic inference with an accent on LLM and agent evaluation, trajectory-level metrics, and adaptive model routing. Focus on designing task suites, measuring cost-quality trade-offs, prototyping routing and escalation policies, and validating methods through traffic replay, shadow testing, and internal pilots.

Location: Singapore, Singapore or Austin, United States

Company

hirify.global is a global technology company providing Bitcoin mining solutions, AI cloud capabilities, and computing infrastructure.

What you will do

  • Build and extend LLM and agent evaluation pipelines for measurable, reliable, and adaptive agentic inference.
  • Develop representative task suites and trajectory-level metrics covering task success, tool use, quality, cost, latency, and token consumption.
  • Research and prototype adaptive model-routing strategies, including cascades, stage-aware routing, uncertainty-aware selection, escalation, and recovery policies.
  • Validate promising methods through offline evaluation, traffic replay, shadow testing, and internal pilots.
  • Collaborate with MaaS and platform engineering teams on production integration while owning evaluation methodology, routing policy, and research prototypes.

Requirements

  • Degree in Computer Science, Machine Learning, Statistics, Electrical Engineering, or a related field, with substantial hands-on experience in LLM evaluation, agentic systems, applied machine learning, or adaptive inference.
  • Strong Python programming skills and practical experience with PyTorch and modern data and evaluation tooling.
  • Experience designing or operating LLM or agent evaluation pipelines, including task- and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis.
  • Implementation-level depth in model selection and routing, uncertainty estimation and calibration, cascading and escalation, or stage-aware agent inference.
  • Experience evaluating multi-turn or tool-using agents and analyzing task completion, tool-call correctness, planning failures, recovery behavior, cost, latency, and token usage.
  • Rigorous experimental practice, including controlled comparisons, statistical analysis, honest baselines, and cost-quality Pareto frontiers.

Nice to have

  • Experience with preference modeling, contextual bandits, online learning, or broader adaptive inference methods.
  • Experience with multi-model APIs, agent harnesses, traffic replay, shadow evaluation, A/B testing, or production model monitoring.
  • Familiarity with tool-calling differences, context windows, prompt caching, reasoning controls, vLLM, or SGLang.
  • Top-tier ML, NLP, or systems publications, or substantial open-source contributions in evaluation, agents, routing, or inference.

Culture & Benefits

  • Inclusive environment valuing authenticity and diverse perspectives.
  • Startup-oriented atmosphere within a fast-growing global company.
  • Opportunity to contribute to new AI, computing, and digital asset projects.
  • Autonomy, personal accountability, rapid growth, and learning opportunities.
  • Training, mentoring, welfare benefits, and professional development support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →