Advanced Workshop: Mastering AI Observability — Doug Guthrie, Braintrust
Summary
This workshop details a comprehensive approach to AI Observability, positioning Braintrust as a platform that enables the creation of a continuous agent improvement 'flywheel.' The core methodology involves capturing massive amounts of agent interaction data (traces) and deriving actionable 'signal' through advanced features like custom scoring, Topics, and automated pattern recognition. This signal is then used to inform development, allowing engineers to automatically generate pull requests (PRs) with proposed code and scorer changes, thereby closing the loop between production performance and development quality.
Key takeaways
-
The Agent Improvement Flywheel
3:40
The goal is to create a closed loop where production data informs development. This involves capturing production traces, deriving insights (signal), making changes, running evaluations (Evals), and repeating the cycle to improve agent quality.
-
Active Observability via Topics
10:00
Beyond traditional failure modes (known unknowns), the Topics feature analyzes traces to identify patterns (e.g., task, sentiment, issues) that users are interacting with, helping uncover 'unknown unknowns' and potential feature requests.
-
Automated Development Workflow
25:00
The Braintrust CLI and coding agents (e.g., Codex) can be used to automate the entire improvement cycle: querying production data via SQL, identifying failure patterns, proposing code changes, and generating PRs for review.
Technical details
-
Observability Foundation
120s
Tracing is fundamental, requiring the ability to parse massive amounts of agent data (spans, tool calls, inputs/outputs) to understand where quality is deteriorating. Braintrust provides the platform for this.
-
Scoring and Evaluation
300s
Signal can be derived using multiple scoring methods: fast, code-based scores (e.g., validating tool call order); LLM judges (for subjective assessment); and LLM judges aligned with human review. Binary scoring is recommended for initial production deployments.
-
Data Pipeline for Topics
700s
The Topics pipeline processes raw traces through a preprocessor, which feeds into a custom facet (e.g., 'support workflow issue'). This summary is then used to create a vector embedding, which finally generates a Topic Map for classification and labeling.
-
Automation and Querying
1000s
Scores and Topics can be wired up to run as automations on incoming logs/traces. The Braintrust CLI allows users to run arbitrary SQL queries against the stored trace data, enabling flexible querying by coding agents.
-
Deployment Architecture
2000s
Braintrust supports a hybrid deployment model, allowing the data plane (traces, data sets, experiments) to reside within a customer's VPC/infrastructure while the control plane remains hosted by Braintrust.
Mentioned resources
- Braintrust CLI
- OpenAI Agents SDK
- Braintrust GitHub Repo
- Braintrust Platform
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.