# How To Make PDFs Your AI Can Actually Read

## Executive summary

The video discusses the critical challenge of PDF accessibility and machine readability, particularly for complex documents like textbooks and scientific papers. Standard PDFs often fail to provide proper structure for screen readers or OCR models, leading to garbled or unusable data. The speaker highlights Data Lab's accessible PDF API, which processes documents into highly structured formats (e.g., separating text, formulas, and alt text) to ensure context and order are preserved. This structured output is valuable not only for users with disabilities but also for improving the reliability of downstream LLM inputs.

## Key takeaways

- The Problem with Standard PDFs: Most PDFs, especially those derived from scans or complex layouts (like math or tables), lack the necessary structural metadata (text layer, proper tagging) required for screen readers or reliable OCR, often resulting in a 'garbled mess' [0:01:30].
- Structured Data Extraction: The accessible PDF API processes documents by identifying and separating structural elements—such as headers, paragraphs, lists, formulas, and alt text—and presenting them in a structured, usable format [0:02:00].
- Utility for AI and Accessibility: Structured PDFs benefit both accessibility (allowing screen readers to read content in the correct order) and AI models. The structured output can be used to feed downstream LLMs, which perform better when given organized data rather than raw images or unstructured text [0:04:20].

## Technical details

- PDF Accessibility Standards: Accessibility requires that a person with a disability (e.g., using a screen reader or Braille display) perceives, operates, and understands the content in the same context and order as a sighted user. This involves correctly tagging elements like formulas and figures [0:00:45].
- Data Structure and Extraction: The API must distinguish between different content types: standard text, formulas (which must be spoken like LaTeX), alt text for figures, and page furniture. This process requires sophisticated OCR and structural analysis beyond basic text extraction [0:01:40].
- LLM Input Optimization: For feeding data into LLMs, structured output is highly beneficial. While the API provides structure, other models (like Shandra) can output different formats, such as HTML or JSON, allowing developers to programmatically handle different data types (e.g., tables vs. lists) [0:04:40].

## Practical implications

- Build engineers can integrate this API into data ingestion pipelines to pre-process complex, non-standard PDF documents before feeding them to LLMs or archival systems.
- The structured output format can be used as a foundational layer for building educational or technical content systems that require high accessibility compliance.
- Developers can utilize the structured output (JSON/HTML) to programmatically differentiate and handle content types (e.g., treating a formula block differently than a paragraph) for downstream applications.

## Topics

PDF Processing, Accessibility (WCAG), OCR, LLMs, Data Structuring, Data Pipelines, Machine Learning, Data Lab Accessible PDF API, O'Reilly book on AI Evals

Source: https://www.youtube.com/watch?v=n1PsY1W1ifc
