Назад
6 дней назад

AI Test Engineer

Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
China
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
AI Test Engineer (AI Agents): Building enterprise testing and evaluation systems for AI applications, knowledge retrieval, tool calls, multi-turn conversations, and workflows with an accent on evaluation datasets, benchmarks, metrics, and regression mechanisms. Focus on designing human and automated evaluation, analyzing logs and call chains, and leading stability, performance, security, and adversarial testing.

AI Test Engineer

Company

01.AI

Conditions

1 week ago

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own testing and evaluation delivery for enterprise Agent projects. You will define test strategies, evaluation plans, and acceptance criteria; identify quality risks; and provide launch, delivery, and customer-acceptance recommendations. You will evaluate AI applications, knowledge retrieval, Agents, tool calls, multi-turn conversations, and workflows. You will build enterprise evaluation systems, datasets, benchmarks, regression mechanisms, Harness capabilities, and evaluation-platform components. You will also lead quality initiatives for stability, performance, security, and adversarial testing.

Requirements

  • 5 or more years of testing, test development, quality, or evaluation experience
  • 2 or more years of AI or Agent evaluation experience
  • Ability to independently deliver complex projects
  • Harness knowledge and practical testing or evaluation Harness experience
  • AI and Agent evaluation methods
  • Evaluation dataset development, metric design, and combined human and automated evaluation
  • Knowledge retrieval, Agent, tool calling, multi-turn conversation, and workflow knowledge
  • Python
  • Enterprise project experience
  • Data isolation, permission management, auditing, release rollback, and private deployment quality requirements
  • Cross-team delivery ability

Responsibilities

  • Define test strategies, evaluation plans, and acceptance criteria for enterprise Agent projects
  • Identify quality risks from customer business and usage scenarios
  • Drive product, algorithm, engineering, and delivery teams to resolve issues
  • Deliver evaluation conclusions and improvement recommendations for launch, delivery, and customer acceptance
  • Test and evaluate AI applications, knowledge retrieval, Agents, tool calls, multi-turn conversations, and workflows
  • Design combined human, automated, and model-assisted evaluation approaches
  • Analyze logs and call chains, drive fixes, and verify regressions
  • Build quality metrics, evaluation datasets, standards, baselines, release thresholds, and feedback mechanisms
  • Build Harness and evaluation-platform capabilities for execution, scoring, comparison, analysis, and reporting
  • Lead stability, performance, security, and adversarial testing initiatives
  • Develop testing and evaluation scripts, tools, or platform components

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -