8 дней назад
Quality Lead, Agentic AI Workflow Evaluation (AI)
75 - 85$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Quality Lead, Agentic AI Workflow Evaluation (AI): Building and operating a quality system for evaluating complex, multi-step agentic AI workflows with an accent on audit design, calibration, rubric health, and reviewer readiness. Focus on identifying systematic error patterns, resolving judgment disagreements, improving evaluation standards, and supporting onsite delivery operations.
Location: In office in San Jose, California, United States
Salary: $75–85 per hour, based on experience, skills, and qualifications.
Company
is a global data engineering company providing data, evaluation frameworks, platforms, and human expertise for generative AI and other AI systems.
What you will do
- Own the engagement quality system, including audit design, sampling strategy, scoring standards, measurement, and reporting.
- Re-score reviewer output, identify recurring error patterns, and provide evidence for performance and coaching conversations.
- Run calibration sessions, resolve scoring disagreements, and document the reasoning for future cases.
- Maintain rubric quality by identifying ambiguous, overlapping, or incomplete criteria and driving improvements with customer quality leads.
- Train and onboard reviewers, define nesting and ramp criteria, and determine production readiness.
- Report quality trends to the Engagement Manager and customer, deputizing for the manager during absences and maintaining required security and facility-access practices.
Requirements
- Bachelor's degree or equivalent practical experience.
- At least 4 years of experience in quality assurance, quality management, or senior review work in annotation, evaluation, trust and safety, or a similarly judgment-intensive domain.
- Direct experience owning a quality function, including designing audits and quality measurement systems.
- Significant experience with AI/ML evaluation, annotation, red-teaming, RLHF, model or agent evaluation, or trust and safety review.
- Hands-on familiarity with agentic systems, tool use, multi-step task execution, sandboxed environments, and common failure modes.
- Strong written communication, calibration experience, reviewer-training experience, and working knowledge of spreadsheets, dashboarding tools, Python, or SQL.
Culture & Benefits
- Work as part of a dedicated onsite team evaluating complex real-world agentic AI workflows for a frontier AI customer.
- Work in isolated test environments focused on safe task completion, user intent, consent, and detailed quality review.
- Partner directly with customer quality leads to build or improve the quality approach.
- Maintain information security, privacy, and facility-access practices required by the onsite customer environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →