Topic

Data Quality

All digests tagged Data Quality

From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss thumbnail

· 21:37

From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss

This talk details the process of scaling an AI-powered chatbot from a small Proof of Concept (POC) to a production system using unstructured data (e.g., 80,000+ SharePoint documents). The core challenge addressed is that AI performance degrades significantly when the underlying data corpus is not curated, containing stale, duplicated, or conflicting information. The solution involves using AI-assisted data curation to generate rich metadata, detect sensitive data (PII/PHI), and maintain data quality (freshness, conflict resolution) to ensure the context provided to Retrieval-Augmented Generation (RAG) agents is accurate and reliable.

Key takeaways

  1. Scaling Challenge: POC vs. Production 5:25

    Initial pilots using only 40 curated files are manageable, but scaling to tens of thousands of files (e.g., 80,000+ SharePoint documents) introduces major data quality hurdles, making the system brittle and difficult to maintain.

  2. Impact of Poor Data Quality 18:30

    If 30% of a data corpus is stale or duplicated, up to 80% of the context an AI agent retrieves can be useless. Maintaining data quality (duplication and freshness) is critical and can double the recall and improve accuracy of multi-hop RAG evaluations.

  3. Automating Data Curation 20:30

    The process can be automated by generating metadata and creating 'data slices'—curated, relevant subsets of data—which can be refreshed on a set schedule to ensure the chatbot never uses stale information.

Watch on YouTube Full article