# Advanced Workshop: Mastering AI Observability — Doug Guthrie, Braintrust

## Executive summary

This workshop details a comprehensive approach to AI Observability, positioning Braintrust as a platform that enables the creation of a continuous agent improvement 'flywheel.' The core methodology involves capturing massive amounts of agent interaction data (traces) and deriving actionable 'signal' through advanced features like custom scoring, Topics, and automated pattern recognition. This signal is then used to inform development, allowing engineers to automatically generate pull requests (PRs) with proposed code and scorer changes, thereby closing the loop between production performance and development quality.

## Key takeaways

- The Agent Improvement Flywheel: The goal is to create a closed loop where production data informs development. This involves capturing production traces, deriving insights (signal), making changes, running evaluations (Evals), and repeating the cycle to improve agent quality.
- Active Observability via Topics: Beyond traditional failure modes (known unknowns), the Topics feature analyzes traces to identify patterns (e.g., task, sentiment, issues) that users are interacting with, helping uncover 'unknown unknowns' and potential feature requests.
- Automated Development Workflow: The Braintrust CLI and coding agents (e.g., Codex) can be used to automate the entire improvement cycle: querying production data via SQL, identifying failure patterns, proposing code changes, and generating PRs for review.

## Technical details

- Observability Foundation: Tracing is fundamental, requiring the ability to parse massive amounts of agent data (spans, tool calls, inputs/outputs) to understand where quality is deteriorating. Braintrust provides the platform for this.
- Scoring and Evaluation: Signal can be derived using multiple scoring methods: fast, code-based scores (e.g., validating tool call order); LLM judges (for subjective assessment); and LLM judges aligned with human review. Binary scoring is recommended for initial production deployments.
- Data Pipeline for Topics: The Topics pipeline processes raw traces through a preprocessor, which feeds into a custom facet (e.g., 'support workflow issue'). This summary is then used to create a vector embedding, which finally generates a Topic Map for classification and labeling.
- Automation and Querying: Scores and Topics can be wired up to run as automations on incoming logs/traces. The Braintrust CLI allows users to run arbitrary SQL queries against the stored trace data, enabling flexible querying by coding agents.
- Deployment Architecture: Braintrust supports a hybrid deployment model, allowing the data plane (traces, data sets, experiments) to reside within a customer's VPC/infrastructure while the control plane remains hosted by Braintrust.

## Practical implications

- Build engineers can integrate Braintrust into CI/CD pipelines to automatically run agent evaluations (Evals) against production data.
- The platform allows for the creation of automated workflows that generate PRs, proposing code or scorer changes based on observed production failures.
- Using the Braintrust CLI enables coding agents to write and execute SQL queries directly against all stored trace data, providing deep, scalable insights.
- The hybrid deployment option ensures that sensitive production data remains within the customer's private cloud infrastructure.

## Topics

AI Observability, Agent Development, MLOps, Build Engineering, LLM Evaluation, Braintrust CLI, OpenAI Agents SDK, Braintrust GitHub Repo, Braintrust Platform

Source: https://www.youtube.com/watch?v=hfEczxdNyvU
