# Why Building an Eval Platform Is Harder Than It Looks — Braintrust

## Executive summary

Building robust agent quality platforms is a complex systems challenge, requiring more than just a UI or spreadsheet. The core goal is maintaining confidence in non-deterministic LLM agents by establishing a continuous improvement loop. This loop requires integrating pre-production evaluation (Evals) with post-production observability. The primary technical hurdle is managing massive, nested JSON agent traces, which necessitates specialized data querying tools like BTQL to support both real-time monitoring and long-running analysis.

## Key takeaways

- Evals vs. Observability: Agent quality relies on two pillars: Evals (testing behavior before production) and Observability (monitoring interactions with real-life users in production). Both are necessary to understand and improve agent performance due to the non-deterministic nature of LLMs.
- The Evolution of Evals Maturity: Eval platforms progress through stages: 1) Spreadsheets (basic logging); 2) Custom UIs (allowing PMs to participate); 3) Side-by-side experiments (tweakable parameters); and 4) Production Flywheels (connecting live failure modes back into offline test cases).
- The Data Challenge of Agent Traces: Agent traces are semi-structured JSON, often containing hundreds of megabytes of data per interaction. This volume and structure overwhelm typical cloud data warehouses, requiring specialized query capabilities for both real-time and long-running analysis.
- Advanced Automation: The most advanced platforms allow coding agents to run evals themselves using natural language prompts (e.g., 'Find me all traces in the last 24 hours where the user had a poor experience. Run the evals for me.'), enabling the surfacing of 'unknown unknowns.'

## Technical details

- Agent Quality Pillars: Agent quality requires both Evals (pre-production testing) and Observability (continuous monitoring in production) because LLMs are non-deterministic, creating inherent variability and risk.
- Eval System Components: A basic eval system requires three components: 1) A way to execute agents against test inputs; 2) A way to view outputs/scores; and 3) A set of test inputs (prompts, requests, or context) to invoke the agent.
- Data Architecture for Tracing: The underlying technology is complex because agent traces are massive, semi-structured JSON payloads. Querying this data requires an abstraction layer, such as BTQL, to handle both real-time ingestion and long-running aggregation queries.
- Improvement Loop Mechanics: The goal is to observe failure modes in production, log every input/output/step of the trace, define failure dimensions, and then use those failure modes to create targeted offline test cases for iteration, thereby preventing regressions.

## Practical implications

- When designing an AI feature, treat the evaluation platform as a core systems component, not just a reporting UI.
- Architect data pipelines to handle high-volume, deeply nested, semi-structured JSON payloads (agent traces) to support both real-time and historical querying.
- Implement a formal 'Improvement Loop' that systematically feeds production failure data back into the offline testing and evaluation process.
- Design the platform to empower non-technical users (PMs, SMEs) to participate in testing and iteration, moving beyond simple logging.

## Topics

AI Agents, LLMs, Data Architecture, Observability, Evaluation Platforms, System Design, Braintrust, Braintrust SQL / BTQL reference

Source: https://www.youtube.com/watch?v=mUQoVz7THu0
