# From Vibes to Production: Evaluating and Shipping AI Agents That Work 201 — Laurie Voss, Arize AI

## Executive summary

The presentation outlines the evolution of observability and evaluation for non-deterministic AI agents, moving beyond manual trace review to automated, continuous improvement loops. The core argument is that as AI applications scale to millions of requests, traditional methods (reading individual traces or even running manual evaluations) become bottlenecks. The solution is 'Signal,' a system that detects patterns and recurring problems across massive datasets of evaluation failures, suggesting automated fixes, generating GitHub issues, or proposing pull requests (PRs) to automatically improve the agent's behavior.

## Key takeaways

- Traces as the Source of Truth: Because AI agents are non-deterministic, the source of truth for agent behavior is not the code, but the traces—which capture every LLM call, tool call, and agent turn. Traces provide visibility into the entire agent workflow, allowing debugging of complex issues like poor search quality or unnecessary turns.
- The Evolution from Traces to Signals: The observability process evolves through stages: 1) Traces (raw data) $\rightarrow$ 2) Evals (LLMs scoring/explaining traces) $\rightarrow$ 3) Signals (automated pattern detection across mass evaluation failures). This shift moves observability from merely reporting what is happening to actively improving the software.
- Automated Improvement Loop (The 2026 Loop): The goal is a self-improving system where the process moves from no observability $\rightarrow$ traces $\rightarrow$ evals $\rightarrow$ signals $\rightarrow$ automated fixes. Signal automates this by continuously monitoring traces and suggesting fixes, which can be implemented as PRs.

## Technical details

- Observability Standards: The system utilizes Open Inference, the open standard used across the observability industry, to track AI application activity. Implementing observability requires only turning on the tracing mechanism, not writing new code.
- Agentic Coding and Skills: The demonstration showed that a coding agent, equipped with pre-installed 'skills' (e.g., `arise-ai` skills), can programmatically pull traces from Arize AX and analyze them for quality issues, even when no explicit error spans exist.
- Signal Functionality: Signal is an agent that continuously monitors traces, detecting patterns in evaluation failures. It can generate GitHub issues, create evaluation datasets, or propose pull requests (PRs) that automatically fix identified problems, allowing the system to 'fix itself.'

## Practical implications

- Build engineers can move from reactive debugging (reading individual traces) to proactive, automated system improvement by implementing Signal.
- The process of defining 'good behavior' can be codified into regression evals, ensuring that automated fixes do not break existing functionality.
- The ability to generate PRs directly from observed failure patterns drastically accelerates the software development life cycle (SDLC) for AI applications.

## Topics

AI Agents, Observability, MLOps, Continuous Integration/Continuous Delivery (CI/CD), LLM Development, Arize AX, Signal, Open Inference, Wonder Toys

Source: https://www.youtube.com/watch?v=F0TNSmbo5hE
