Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex
This talk details the evolution of document search, contrasting the simple, text-based nature of codebases (where plain file traversal and `grep` suffice) with the complexity of enterprise data (multimodal PDFs, schematics, and slides). The solution presented is an agentic 'harness' that allows an AI agent to dynamically choose the best retrieval method—from hybrid search to metadata filtering and specialized document reading—to ground answers from massive, messy, and permissioned corporate data sets. Key components include hybrid retrieval, structured parsing via LlamaParse/LiteParse, and managing production concerns like multi-tenancy and data freshness.
Key takeaways
-
Code vs. Company Data Search
5:27
Plain file traversal (like Claude Code's approach) works well for code because it is small, text-based, and structured. However, company data (PDFs, schematics, etc.) is huge, multimodal, and lacks a simple folder hierarchy, necessitating advanced retrieval methods.
-
The Agentic Search Harness
13:42
Instead of relying on a single method, agents should be given a suite of tools (a 'harness') to choose from: hybrid search (semantic + keyword), metadata filtering, file listing/traversal, and specialized document reading.
-
Handling Complex Documents
22:07
For multimodal documents, specialized parsing is required. Tools like LlamaParse and LiteParse are recommended to extract structured content (tables, charts) and generate page screenshots, which aids in spatial reasoning and improves Retrieval-Augmented Generation (RAG).
-
Production Scaling Concerns
20:42
When deploying in production, engineers must address multi-tenancy, granular permissioning (e.g., pulling metadata from Google Drive/SharePoint), and data freshness/synchronization, which are major challenges for pre-indexing systems.