I Ranked Every AI Eval Concept From Worst to Best
This video provides a comprehensive ranking of 28 common AI evaluation (eval) practices, grading them from S-tier (best) to F-tier (worst). The central thesis is that effective evaluation must be data-driven: teams should look at their actual product data first to determine which evals are necessary. The speaker strongly advises against using generic metrics, off-the-shelf tools, or relying on top-down brainstorming, emphasizing instead the use of custom annotation tools and domain expertise.
Key takeaways
-
Prioritize Data Analysis Over Brainstorming
The most valuable activity is looking at the data itself. Writing evals before analyzing real-world failures (a top-down approach) is inefficient. Error analysis (S-tier) is crucial for identifying what is actually broken in the application.
-
Domain Expertise is Non-Negotiable
Domain experts (e.g., a lawyer for a legal AI) must write the prompts and annotate the data, not developers. This prevents a separation of concerns and ensures the evaluation criteria are accurate.
-
Favor Code-Based and Binary Checks
11:57
Code-based checks are the cheapest and most effective form of evaluation. Furthermore, using binary pass/fail metrics is preferred over subjective art scales (like 1 to 5) because binary outcomes cut through noise and clarify decision-making.
-
Validate LLM Judges with Human Labels
To trust an LLM judge, it must be validated against human labels. This process is essential because generic or off-the-shelf prompts rarely make sense for a specific product.