From Vibes to Production: Evaluating and Shipping AI Agents That Work 201 — Laurie Voss, Arize AI
The presentation outlines the evolution of observability and evaluation for non-deterministic AI agents, moving beyond manual trace review to automated, continuous improvement loops. The core argument is that as AI applications scale to millions of requests, traditional methods (reading individual traces or even running manual evaluations) become bottlenecks. The solution is 'Signal,' a system that detects patterns and recurring problems across massive datasets of evaluation failures, suggesting automated fixes, generating GitHub issues, or proposing pull requests (PRs) to automatically improve the agent's behavior.
Key takeaways
-
Traces as the Source of Truth
0:03
Because AI agents are non-deterministic, the source of truth for agent behavior is not the code, but the traces—which capture every LLM call, tool call, and agent turn. Traces provide visibility into the entire agent workflow, allowing debugging of complex issues like poor search quality or unnecessary turns.
-
The Evolution from Traces to Signals
0:09
The observability process evolves through stages: 1) Traces (raw data) $\rightarrow$ 2) Evals (LLMs scoring/explaining traces) $\rightarrow$ 3) Signals (automated pattern detection across mass evaluation failures). This shift moves observability from merely reporting what is happening to actively improving the software.
-
Automated Improvement Loop (The 2026 Loop)
0:16
The goal is a self-improving system where the process moves from no observability $\rightarrow$ traces $\rightarrow$ evals $\rightarrow$ signals $\rightarrow$ automated fixes. Signal automates this by continuously monitoring traces and suggesting fixes, which can be implemented as PRs.