# Evals in AI: A Deep Dive — Tejas Kumar, IBM

## Executive summary

This deep dive explains why and how AI evaluation (Evals) is critical for building robust, reliable AI applications. Evals serve as 'fuzzy unit tests' for non-deterministic systems, providing 'ahead of time' reliability that complements 'just in time' harnesses. The process involves creating structured evaluation suites—including test cases, expected behavior, and a judge—and integrating them into CI gates to prevent deployment of insecure or unreliable models. The speaker emphasizes that Evals must be treated as a living dataset, requiring continuous refinement using techniques like Retrieval-Augmented Generation (RAG) to incorporate up-to-date policies and context.

## Key takeaways

- Evals provide Ahead-of-Time Reliability: While harnesses provide 'just in time' reliability (e.g., a user prompting the agent), Evals allow for proactive, 'ahead of time' testing, making them a crucial layer for system robustness. The combination of Evals and harnesses provides comprehensive reliability.
- Evals Mitigate Security Exploits: A solid set of Evals can quickly identify potential security vulnerabilities, preventing dangerous actions (like unauthorized account mutations) before deployment, which is a key benefit over relying solely on runtime harnesses.
- The Eval Architecture: A complete Eval suite requires four components: 1) Test cases (inputs/scenarios), 2) Expected behavior (desired outcomes), 3) A Judge (to determine acceptability), and 4) An aggregation method (to produce a probability score, e.g., 80% agreement).
- Addressing Judge Biases: LLM judges are prone to biases, including Position Bias (favoring the first option), Sycophancy (favoring the nicest response), Self-Reference (favoring text from the same model family), and Verbosity (favoring longer answers). These biases must be accounted for in the evaluation design.
- Productionizing Evals with CI Gates: The safest deployment method is to implement the Eval suite as a CI gate. The system should only pass if the judge's agreement with the human-defined 'golden dataset' meets a predefined threshold (e.g., 80%).

## Technical details

- Eval vs. Unit Tests: Unit tests are deterministic, asserting that a function returns an exact number (e.g., `add(1, 2)` must equal 3). Evals are designed for non-deterministic systems, asserting that an agent behaves acceptably across many scenarios, making them 'fuzzy unit tests.'
- Evaluation Techniques: Evaluation techniques include: 1) Exact Match (checking tool calls/arguments), 2) Schema Validation (validating JSON blobs/message envelopes), 3) Pair-wise Comparison (comparing two options), and 4) Pointwise Comparison (asking a single yes/no question). Pointwise comparison is generally preferred over pair-wise due to its higher reliability.
- Training the Judge: To train an LLM judge, one must create a 'golden dataset' (a large collection of Question/Answer/Human Verdict). The process involves removing the human score, sending the Q/A pair to the judge, and calculating the delta (agreement percentage) against the human verdict. The goal is typically 80-85% agreement.
- Maintaining Policy Context (RAG): To ensure the judge and the production agent share the same knowledge base, the policy must be retrieved dynamically. Using Retrieval-Augmented Generation (RAG) to inject the actual, up-to-date policy into the prompt is crucial for accurate evaluation.

## Practical implications

- Implement Evals as a mandatory CI gate, failing the build if the judge's agreement falls below the established threshold (e.g., 80%).
- Treat the Eval dataset as a living asset. When in production, continuously collect and manually score real-world traffic data to keep the judge relevant.
- When designing the judge, ensure the judge model family is different from the model family generating the answers to mitigate self-reference bias.
- Always incorporate the authoritative, current policy document into the prompt context (via RAG) to ground the judge's decisions.

## Topics

AI Evaluation (Evals), LLM Judging, Reliability Engineering, CI/CD Pipelines, RAG Architecture, Non-Deterministic Systems, OpenRAG, Dockling

Source: https://www.youtube.com/watch?v=NdrSPm6NCdk
