AI Engineer

Advanced Workshop: Mastering AI Observability — Doug Guthrie, Braintrust

Published 2026-10-06 · Duration 1:51:37

Summary

This workshop details a comprehensive approach to AI Observability, positioning Braintrust as a platform that enables the creation of a continuous agent improvement 'flywheel.' The core methodology involves capturing massive amounts of agent interaction data (traces) and deriving actionable 'signal' through advanced features like custom scoring, Topics, and automated pattern recognition. This signal is then used to inform development, allowing engineers to automatically generate pull requests (PRs) with proposed code and scorer changes, thereby closing the loop between production performance and development quality.

Download summary

Key takeaways

  1. The Agent Improvement Flywheel 3:40

    The goal is to create a closed loop where production data informs development. This involves capturing production traces, deriving insights (signal), making changes, running evaluations (Evals), and repeating the cycle to improve agent quality.

  2. Active Observability via Topics 10:00

    Beyond traditional failure modes (known unknowns), the Topics feature analyzes traces to identify patterns (e.g., task, sentiment, issues) that users are interacting with, helping uncover 'unknown unknowns' and potential feature requests.

  3. Automated Development Workflow 25:00

    The Braintrust CLI and coding agents (e.g., Codex) can be used to automate the entire improvement cycle: querying production data via SQL, identifying failure patterns, proposing code changes, and generating PRs for review.

Technical details

  • Observability Foundation 120s

    Tracing is fundamental, requiring the ability to parse massive amounts of agent data (spans, tool calls, inputs/outputs) to understand where quality is deteriorating. Braintrust provides the platform for this.

  • Scoring and Evaluation 300s

    Signal can be derived using multiple scoring methods: fast, code-based scores (e.g., validating tool call order); LLM judges (for subjective assessment); and LLM judges aligned with human review. Binary scoring is recommended for initial production deployments.

  • Data Pipeline for Topics 700s

    The Topics pipeline processes raw traces through a preprocessor, which feeds into a custom facet (e.g., 'support workflow issue'). This summary is then used to create a vector embedding, which finally generates a Topic Map for classification and labeling.

  • Automation and Querying 1000s

    Scores and Topics can be wired up to run as automations on incoming logs/traces. The Braintrust CLI allows users to run arbitrary SQL queries against the stored trace data, enabling flexible querying by coding agents.

  • Deployment Architecture 2000s

    Braintrust supports a hybrid deployment model, allowing the data plane (traces, data sets, experiments) to reside within a customer's VPC/infrastructure while the control plane remains hosted by Braintrust.

Mentioned resources

  • Braintrust CLI (Command Line Tool)
  • OpenAI Agents SDK (Software Development Kit)
  • Braintrust GitHub Repo (Code Repository)
  • Braintrust Platform (SaaS/Observability Platform)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.