# 8 Claude Code skills I use for building better AI evals

## Executive summary

This technical deep dive introduces eight specialized Claude Code skills designed to improve the robustness of AI product evaluations (AI Evals). The core message is that teams must analyze real product failures before designing evaluation metrics. The two most critical skills are 'Error Discovery,' which guides the process of finding failure modes by reviewing data traces, and 'Eval Audit,' which systematically checks the entire evaluation setup against common industry mistakes.

## Key takeaways

- Prioritize Failure Analysis Over Planning: The biggest mistake in AI product development is writing evaluation plans before analyzing actual failures observed in the product. The speaker recommends starting with the Error Discovery and Eval Audit skills.

## Technical details

- Error Discovery Skill: This skill guides the agent through multiple phases to find failure modes: 1) Data understanding (parsing structure and identifying dimensions of variation). 2) Annotation Interface Design (using principles like Gestalt). 3) Building the review interface. It intelligently samples data, prioritizing diversity (breadth) over frequency, and utilizes an interactive review loop that updates sampling as new errors are found.
- Eval Audit Skill: This skill performs diagnostic checks on the evaluation pipeline to prevent common mistakes. Checks include ensuring error analysis was performed, confirming failure categories are observed in data, mandating binary pass/fail metrics (avoiding Likert scales), validating LLM judges against human labels, and checking for proper data splitting to prevent overfitting.
- LLM Judge Validation: To ensure an LLM judge is reliable, the Eval Audit skill mandates checking it against human labels. It also emphasizes proper data hygiene and splitting to prevent the judge from 'cheating' during evaluation.
- Synthetic Data Generation: The 'Generate Synthetic Data' skill is useful for pre-product stress testing when real user data is unavailable, helping to maximize the diversity of the synthetic data.

## Practical implications

- Integrate systematic error analysis (Error Discovery) into the early stages of the product lifecycle, rather than waiting for formal evaluation.
- Use the Eval Audit skill as a mandatory pre-evaluation checklist to ensure the evaluation setup is robust and avoids common pitfalls (e.g., using non-binary metrics, neglecting judge validation).
- Focus on maximizing data diversity when sampling traces to ensure comprehensive coverage of failure modes.
- Leverage the skills to build automated, structured review interfaces for annotating complex, multi-turn conversational data.

## Topics

AI Evals, LLM Evaluation, Claude Code Skills, RAG Pipelines, Data Annotation, Failure Mode Analysis, The eval skills on GitHub, AI evals course flashcards

Source: https://www.youtube.com/watch?v=c5Ur73GosR4
