Advanced Workshop: Mastering AI Observability — Doug Guthrie, Braintrust
This workshop details a comprehensive approach to AI Observability, positioning Braintrust as a platform that enables the creation of a continuous agent improvement 'flywheel.' The core methodology involves capturing massive amounts of agent interaction data (traces) and deriving actionable 'signal' through advanced features like custom scoring, Topics, and automated pattern recognition. This signal is then used to inform development, allowing engineers to automatically generate pull requests (PRs) with proposed code and scorer changes, thereby closing the loop between production performance and development quality.
Key takeaways
-
The Agent Improvement Flywheel
3:40
The goal is to create a closed loop where production data informs development. This involves capturing production traces, deriving insights (signal), making changes, running evaluations (Evals), and repeating the cycle to improve agent quality.
-
Active Observability via Topics
10:00
Beyond traditional failure modes (known unknowns), the Topics feature analyzes traces to identify patterns (e.g., task, sentiment, issues) that users are interacting with, helping uncover 'unknown unknowns' and potential feature requests.
-
Automated Development Workflow
25:00
The Braintrust CLI and coding agents (e.g., Codex) can be used to automate the entire improvement cycle: querying production data via SQL, identifying failure patterns, proposing code changes, and generating PRs for review.