# How to Evaluate OCR for AI Agents

## Executive summary

This technical deep dive addresses the critical failure point in AI agent pipelines: the quality of input data derived from complex documents. The speaker, Isaac Flath, argues that when agents fail, the error is often not a model hallucination, but an upstream data issue originating from OCR inaccuracies, structural misinterpretations, or missing contextual information (like watermarks or draft status). He details the necessity of specialized annotation apps and rigorous evaluation frameworks that test the entire pipeline—OCR, retrieval, and the agent loop—to pinpoint the true root cause of failure.

## Key takeaways

- Document Complexity vs. Simple PDFs: Documents are defined as anything that stores information but is not a database (e.g., scanned forms, contracts). This includes complex structures like tables and grouped fields, which are much harder to process than simple PDFs.
- The Three Parts of Document Evaluation: Evaluating document-based AI requires testing three distinct components: the OCR process, the retrieval mechanism, and the final agent loop. Failure in any one area can lead to an incorrect final answer.
- Root Cause Analysis (RCA) Workflow: When an agent fails, the process involves working backward from the wrong answer to determine if the error is due to model hallucination, context window repetition, or, most commonly, an extraction error from the source document.
- Annotation App Necessity: Specialized annotation apps are critical for error analysis, allowing users to tie bounding boxes and notes directly to specific locations on the original PDF, enabling precise identification of extraction failures.

## Technical details

- OCR Model Comparison: The speaker compares traditional ML pipelines (detect line -> predict text -> predict layout) with modern vision models (e.g., Shandra, Gemini Flash), which predict output in a single pass by treating the page like a PNG image.
- Document Data Types: PDFs can be complex, containing multiple layers: scanned images (no text layer), text-based forms (fillable), or a mix of both. This complexity increases the difficulty of accurate extraction.
- Evaluation Techniques: Advanced testing methods include running simple arithmetic checks (e.g., do the numbers add up to the total?) to validate data integrity, providing a fast test without rerunning the entire agent loop.
- Agent Tooling and Pipelines: The process involves configuring multiple components: OCR model selection, retrieval approach (semantic search, SQL queries), and the agent harness (e.g., Prime Agent, custom tools).

## Practical implications

- Build robust data pipelines that treat OCR output as a potential failure point, rather than a reliable source of truth.
- Implement specialized annotation tools that allow engineers to perform granular, localized error analysis (e.g., bounding box annotation) to pinpoint data extraction failures.
- Design evaluation tests that validate data consistency (e.g., summing values) rather than just testing the final answer.
- When building agents, always consider the state of the document (e.g., 'Draft,' 'Confidential,' 'Example') and ensure this metadata is extracted and used in the prompt/agent logic.

## Topics

AI Agents, OCR, Document Processing, Data Pipeline Engineering, Natural Language Processing, How to choose an OCR model, Document answers you can check, AI Evals October 2026 cohort

Source: https://www.youtube.com/watch?v=UPEzmVR0dJ8
