AI Engineer

The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI

Published 2026-07-24 · Duration 6:06

Summary

The complexity of modern AI agents—which now incorporate reasoning, tool calls, and long multi-step loops—has rendered traditional evaluation methods insufficient. The talk argues that the future of evaluating these systems lies in moving from static deterministic checks or fixed LLM rubrics to 'Agent as a Judge,' which performs adaptive dynamic analysis to uncover subtle failure modes.

Download summary

Key takeaways

  1. Evals are critical for AI maturity

    Evals have become essential for serious AI teams, with the industry noting that they catch all failures and fuel continual learning loops. Arize reports running over 100 million evals monthly.

  2. Agent complexity breaks traditional evals 3:26

    As agents evolved from simple prompt answering (2023) to complex, multi-step loops with sub-agents and dynamic UI creation, the failure modes became fundamentally different, exceeding the scope of classical LLM as a Judge checks.

  3. The future requires adaptive evaluation 5:45

    While deterministic checks and LLM-as-a-Judge are valuable, the next step is 'Agent as a Judge,' which provides adaptive dynamic analysis to find failure modes that were previously undetectable.

Technical details

  • Evaluation Evolution 345s

    Evals have progressed through three stages: 1) Deterministic checks (defining what can be defined up front); 2) LLM as a Judge (adding analysis beyond fixed rules, using a fixed rubric/score); and 3) Agent as a Judge (performing adaptive dynamic analysis to hunt for unforeseen failure modes).

  • Agent Capabilities & Failure Modes 206s

    Modern agents can perform long-horizon tasks, create new UIs based on user interaction (creating a fundamentally different trajectory), and utilize tool calls and deep reasoning. Failures often involve inefficient loops or context loss.

  • Signal Tooling 345s

    Arize introduced 'Signal,' a long-running agent designed to read traces, discover patterns of issues, and identify subtle failures (e.g., repeated tool calls or inefficient loops) that classical LLM evals would miss.

Mentioned resources

  • Arize AI / Signal (Tooling/Product Demo)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.