10 часов назад
Applied Data Science Role (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Applied Data Science Role (AI): Building evaluation pipelines, datasets, graders, and metrics to measure and improve the quality of AI agents handling real-world tasks with an accent on statistical rigor, production systems, and multi-step agent behavior. Focus on calibrating LLM-as-a-judge systems, analyzing traces and tool calls, identifying failure causes, and validating improvements without unacceptable regressions.
Location: On-site in Palo Alto, California, United States; the listing also identifies the location type as remote.
Company
develops AI agents for real-world tasks such as scheduling, email, browser automation, and business software workflows.
What you will do
- Architect and maintain automated evaluation pipelines across agent capabilities and product surfaces.
- Define success criteria, gold datasets, regression suites, and metrics for task success, tool selection, instruction adherence, factual consistency, latency, cost, and reliability.
- Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and measure false positives, false negatives, variance, and grader agreement.
- Analyze traces, tool calls, and production outcomes to identify root causes and develop failure taxonomies.
- Run offline experiments and use production evidence to compare models, prompts, and capability implementations.
- Build dashboards and reports, and partner with engineering to verify quality improvements without unacceptable regressions.
Requirements
- 5+ years of experience in data science, machine learning, or analytics, focused on evaluation systems, metrics frameworks, or production quality measurement.
- Experience designing and implementing evaluation and grading frameworks for production ML or AI products.
- Production-quality Python and SQL skills for automated pipelines and analysis at scale.
- Strong knowledge of evaluation methodology, dataset construction, metric selection, and statistical and experimental design.
- Experience creating ground-truth data, including labeling guidelines, annotation quality control, ambiguity resolution, and dataset maintenance.
- Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes.
Nice to have
- Experience with agentic systems, multi-step task evaluation, or consumer-facing production ML.
Culture & Benefits
- Work at the intersection of evaluation design, statistics, and production systems.
- Collaborate with engineering, product, and leadership to make evaluation results clear and actionable.
- Visa sponsorship is not available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →