From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss
Summary
This talk details the process of scaling an AI-powered chatbot from a small Proof of Concept (POC) to a production system using unstructured data (e.g., 80,000+ SharePoint documents). The core challenge addressed is that AI performance degrades significantly when the underlying data corpus is not curated, containing stale, duplicated, or conflicting information. The solution involves using AI-assisted data curation to generate rich metadata, detect sensitive data (PII/PHI), and maintain data quality (freshness, conflict resolution) to ensure the context provided to Retrieval-Augmented Generation (RAG) agents is accurate and reliable.
Key takeaways
-
Scaling Challenge: POC vs. Production
5:25
Initial pilots using only 40 curated files are manageable, but scaling to tens of thousands of files (e.g., 80,000+ SharePoint documents) introduces major data quality hurdles, making the system brittle and difficult to maintain.
-
Impact of Poor Data Quality
18:30
If 30% of a data corpus is stale or duplicated, up to 80% of the context an AI agent retrieves can be useless. Maintaining data quality (duplication and freshness) is critical and can double the recall and improve accuracy of multi-hop RAG evaluations.
-
Automating Data Curation
20:30
The process can be automated by generating metadata and creating 'data slices'—curated, relevant subsets of data—which can be refreshed on a set schedule to ensure the chatbot never uses stale information.
Technical details
-
AI-Assisted Taxonomy and Metadata Tagging
635s
The system allows users to define a tagging tree (taxonomy) and leverage AI to auto-suggest relevant tags for documents. Metadata generation includes 'evidence' (showing where in the document the AI found the evidence) and a 'confidence score' for verification.
-
Data Quality Dimensions
425s
Beyond standard data quality (completeness, conformity, accuracy), document-specific quality checks include identifying duplicate information, conflicting information, and data freshness (ensuring files are within a specific date range).
-
Sensitive Data Detection
1004s
The platform can scan documents for sensitive data types, including PII and PHI, using both AI-based and pattern-based detection methods, allowing for filtering before the data reaches the chatbot.
-
Context Repository Generation
1180s
The SDK can be used to curate context files for coding agents (like CodeX). Instead of having the agent search individual files, it looks at a curated context repository that defines the purpose, document types, and key topics of entire folders.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.