AI Engineer

Evals in AI: A Deep Dive — Tejas Kumar, IBM

Published 2026-10-05 · Duration 59:49

Summary

This deep dive explains why and how AI evaluation (Evals) is critical for building robust, reliable AI applications. Evals serve as 'fuzzy unit tests' for non-deterministic systems, providing 'ahead of time' reliability that complements 'just in time' harnesses. The process involves creating structured evaluation suites—including test cases, expected behavior, and a judge—and integrating them into CI gates to prevent deployment of insecure or unreliable models. The speaker emphasizes that Evals must be treated as a living dataset, requiring continuous refinement using techniques like Retrieval-Augmented Generation (RAG) to incorporate up-to-date policies and context.

Download summary

Key takeaways

  1. Evals provide Ahead-of-Time Reliability 3:33

    While harnesses provide 'just in time' reliability (e.g., a user prompting the agent), Evals allow for proactive, 'ahead of time' testing, making them a crucial layer for system robustness. The combination of Evals and harnesses provides comprehensive reliability.

  2. Evals Mitigate Security Exploits 10:52

    A solid set of Evals can quickly identify potential security vulnerabilities, preventing dangerous actions (like unauthorized account mutations) before deployment, which is a key benefit over relying solely on runtime harnesses.

  3. The Eval Architecture 17:03

    A complete Eval suite requires four components: 1) Test cases (inputs/scenarios), 2) Expected behavior (desired outcomes), 3) A Judge (to determine acceptability), and 4) An aggregation method (to produce a probability score, e.g., 80% agreement).

  4. Addressing Judge Biases 25:21

    LLM judges are prone to biases, including Position Bias (favoring the first option), Sycophancy (favoring the nicest response), Self-Reference (favoring text from the same model family), and Verbosity (favoring longer answers). These biases must be accounted for in the evaluation design.

  5. Productionizing Evals with CI Gates 42:18

    The safest deployment method is to implement the Eval suite as a CI gate. The system should only pass if the judge's agreement with the human-defined 'golden dataset' meets a predefined threshold (e.g., 80%).

Technical details

  • Eval vs. Unit Tests 823s

    Unit tests are deterministic, asserting that a function returns an exact number (e.g., `add(1, 2)` must equal 3). Evals are designed for non-deterministic systems, asserting that an agent behaves acceptably across many scenarios, making them 'fuzzy unit tests.'

  • Evaluation Techniques 1245s

    Evaluation techniques include: 1) Exact Match (checking tool calls/arguments), 2) Schema Validation (validating JSON blobs/message envelopes), 3) Pair-wise Comparison (comparing two options), and 4) Pointwise Comparison (asking a single yes/no question). Pointwise comparison is generally preferred over pair-wise due to its higher reliability.

  • Training the Judge 2700s

    To train an LLM judge, one must create a 'golden dataset' (a large collection of Question/Answer/Human Verdict). The process involves removing the human score, sending the Q/A pair to the judge, and calculating the delta (agreement percentage) against the human verdict. The goal is typically 80-85% agreement.

  • Maintaining Policy Context (RAG) 3000s

    To ensure the judge and the production agent share the same knowledge base, the policy must be retrieved dynamically. Using Retrieval-Augmented Generation (RAG) to inject the actual, up-to-date policy into the prompt is crucial for accurate evaluation.

Mentioned resources

  • OpenRAG (Framework/Tool)
  • Dockling (Research Project)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.