AI Engineer

From Vibes to Production: Evaluating and Shipping AI Agents That Work 201 — Laurie Voss, Arize AI

Published 2026-10-05 · Duration 42:17

Summary

The presentation outlines the evolution of observability and evaluation for non-deterministic AI agents, moving beyond manual trace review to automated, continuous improvement loops. The core argument is that as AI applications scale to millions of requests, traditional methods (reading individual traces or even running manual evaluations) become bottlenecks. The solution is 'Signal,' a system that detects patterns and recurring problems across massive datasets of evaluation failures, suggesting automated fixes, generating GitHub issues, or proposing pull requests (PRs) to automatically improve the agent's behavior.

Download summary

Key takeaways

  1. Traces as the Source of Truth 0:03

    Because AI agents are non-deterministic, the source of truth for agent behavior is not the code, but the traces—which capture every LLM call, tool call, and agent turn. Traces provide visibility into the entire agent workflow, allowing debugging of complex issues like poor search quality or unnecessary turns.

  2. The Evolution from Traces to Signals 0:09

    The observability process evolves through stages: 1) Traces (raw data) $\rightarrow$ 2) Evals (LLMs scoring/explaining traces) $\rightarrow$ 3) Signals (automated pattern detection across mass evaluation failures). This shift moves observability from merely reporting what is happening to actively improving the software.

  3. Automated Improvement Loop (The 2026 Loop) 0:16

    The goal is a self-improving system where the process moves from no observability $\rightarrow$ traces $\rightarrow$ evals $\rightarrow$ signals $\rightarrow$ automated fixes. Signal automates this by continuously monitoring traces and suggesting fixes, which can be implemented as PRs.

Technical details

  • Observability Standards 16s

    The system utilizes Open Inference, the open standard used across the observability industry, to track AI application activity. Implementing observability requires only turning on the tracing mechanism, not writing new code.

  • Agentic Coding and Skills 20s

    The demonstration showed that a coding agent, equipped with pre-installed 'skills' (e.g., `arise-ai` skills), can programmatically pull traces from Arize AX and analyze them for quality issues, even when no explicit error spans exist.

  • Signal Functionality 22s

    Signal is an agent that continuously monitors traces, detecting patterns in evaluation failures. It can generate GitHub issues, create evaluation datasets, or propose pull requests (PRs) that automatically fix identified problems, allowing the system to 'fix itself.'

Mentioned resources

  • Arize AX (Observability Platform)
  • Signal (AI Agent/Feature)
  • Open Inference (Open Standard)
  • Wonder Toys (Demonstration Application)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.