# From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss

## Executive summary

This talk details the process of scaling an AI-powered chatbot from a small Proof of Concept (POC) to a production system using unstructured data (e.g., 80,000+ SharePoint documents). The core challenge addressed is that AI performance degrades significantly when the underlying data corpus is not curated, containing stale, duplicated, or conflicting information. The solution involves using AI-assisted data curation to generate rich metadata, detect sensitive data (PII/PHI), and maintain data quality (freshness, conflict resolution) to ensure the context provided to Retrieval-Augmented Generation (RAG) agents is accurate and reliable.

## Key takeaways

- Scaling Challenge: POC vs. Production: Initial pilots using only 40 curated files are manageable, but scaling to tens of thousands of files (e.g., 80,000+ SharePoint documents) introduces major data quality hurdles, making the system brittle and difficult to maintain.
- Impact of Poor Data Quality: If 30% of a data corpus is stale or duplicated, up to 80% of the context an AI agent retrieves can be useless. Maintaining data quality (duplication and freshness) is critical and can double the recall and improve accuracy of multi-hop RAG evaluations.
- Automating Data Curation: The process can be automated by generating metadata and creating 'data slices'—curated, relevant subsets of data—which can be refreshed on a set schedule to ensure the chatbot never uses stale information.

## Technical details

- AI-Assisted Taxonomy and Metadata Tagging: The system allows users to define a tagging tree (taxonomy) and leverage AI to auto-suggest relevant tags for documents. Metadata generation includes 'evidence' (showing where in the document the AI found the evidence) and a 'confidence score' for verification.
- Data Quality Dimensions: Beyond standard data quality (completeness, conformity, accuracy), document-specific quality checks include identifying duplicate information, conflicting information, and data freshness (ensuring files are within a specific date range).
- Sensitive Data Detection: The platform can scan documents for sensitive data types, including PII and PHI, using both AI-based and pattern-based detection methods, allowing for filtering before the data reaches the chatbot.
- Context Repository Generation: The SDK can be used to curate context files for coding agents (like CodeX). Instead of having the agent search individual files, it looks at a curated context repository that defines the purpose, document types, and key topics of entire folders.

## Practical implications

- Build engineers can significantly reduce data preparation time (from months to days) by automating data curation and metadata enrichment.
- The ability to generate context files for coding agents improves the context provided to LLMs, potentially reducing token costs and improving agent performance.
- Implementing automated data slice workflows ensures that AI systems remain current and reliable by automatically refreshing the data source when new files are added or deleted.
- Mitigating compliance risk by proactively identifying and filtering sensitive data (PII/PHI) before it enters the data store.

## Topics

AI, Data Governance, RAG, Unstructured Data, Metadata, Data Quality, SharePoint, Collibra unstructured data for AI, Collibra acquires Deasy Labs, Deasy Labs, Collibra

Source: https://www.youtube.com/watch?v=wzWNYDY7toc
