# I Ranked Every AI Eval Concept From Worst to Best

## Executive summary

This video provides a comprehensive ranking of 28 common AI evaluation (eval) practices, grading them from S-tier (best) to F-tier (worst). The central thesis is that effective evaluation must be data-driven: teams should look at their actual product data first to determine which evals are necessary. The speaker strongly advises against using generic metrics, off-the-shelf tools, or relying on top-down brainstorming, emphasizing instead the use of custom annotation tools and domain expertise.

## Key takeaways

- Prioritize Data Analysis Over Brainstorming: The most valuable activity is looking at the data itself. Writing evals before analyzing real-world failures (a top-down approach) is inefficient. Error analysis (S-tier) is crucial for identifying what is actually broken in the application.
- Domain Expertise is Non-Negotiable: Domain experts (e.g., a lawyer for a legal AI) must write the prompts and annotate the data, not developers. This prevents a separation of concerns and ensures the evaluation criteria are accurate.
- Favor Code-Based and Binary Checks: Code-based checks are the cheapest and most effective form of evaluation. Furthermore, using binary pass/fail metrics is preferred over subjective art scales (like 1 to 5) because binary outcomes cut through noise and clarify decision-making.
- Validate LLM Judges with Human Labels: To trust an LLM judge, it must be validated against human labels. This process is essential because generic or off-the-shelf prompts rarely make sense for a specific product.

## Technical details

- S-Tier Practices (Recommended): Implement error analysis (0:51) and build custom annotation tools (2:03) to remove friction when reviewing traces. Use code-based checks (7:17) as they are the cheapest form of evaluation. Finally, validate LLM judges against human labels (10:54) and ensure domain experts write the prompts (11:59).
- F-Tier Practices (Avoid): Avoid generic metrics (e.g., helpfulness score) (0:27), off-the-shelf LLM judges (3:25), or using public benchmarks (5:24) as substitutes for product-specific evals. These are often generic and fail to correlate with product-specific failures.
- Annotation and Prompting: The speaker advises appointing a 'benevolent dictator' (8:24) for data annotation to minimize cost and complexity. When designing the application, focus on 'Designing for verification' (8:24) by showing intermediate outputs, not just the final answer.

## Practical implications

- Shift focus from building evaluation suites to analyzing failure data (error analysis).
- Invest in custom tooling (e.g., using Claude or Codeex) to streamline the process of reviewing and annotating traces.
- When designing testing pipelines, prioritize the use of cheap, deterministic checks (like code-based tests) over expensive LLM-based judges.
- Ensure that the people defining the evaluation criteria are subject matter experts, not general developers.

## Topics

AI Evaluation (Evals), LLM Judging, Data Annotation, Build Process Optimization, Software Testing Methodologies, AI Evals Guide

Source: https://www.youtube.com/watch?v=Bc15F1xlbTA
