Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex
Summary
This talk details the evolution of document search, contrasting the simple, text-based nature of codebases (where plain file traversal and `grep` suffice) with the complexity of enterprise data (multimodal PDFs, schematics, and slides). The solution presented is an agentic 'harness' that allows an AI agent to dynamically choose the best retrieval method—from hybrid search to metadata filtering and specialized document reading—to ground answers from massive, messy, and permissioned corporate data sets. Key components include hybrid retrieval, structured parsing via LlamaParse/LiteParse, and managing production concerns like multi-tenancy and data freshness.
Key takeaways
-
Code vs. Company Data Search
5:27
Plain file traversal (like Claude Code's approach) works well for code because it is small, text-based, and structured. However, company data (PDFs, schematics, etc.) is huge, multimodal, and lacks a simple folder hierarchy, necessitating advanced retrieval methods.
-
The Agentic Search Harness
13:42
Instead of relying on a single method, agents should be given a suite of tools (a 'harness') to choose from: hybrid search (semantic + keyword), metadata filtering, file listing/traversal, and specialized document reading.
-
Handling Complex Documents
22:07
For multimodal documents, specialized parsing is required. Tools like LlamaParse and LiteParse are recommended to extract structured content (tables, charts) and generate page screenshots, which aids in spatial reasoning and improves Retrieval-Augmented Generation (RAG).
-
Production Scaling Concerns
20:42
When deploying in production, engineers must address multi-tenancy, granular permissioning (e.g., pulling metadata from Google Drive/SharePoint), and data freshness/synchronization, which are major challenges for pre-indexing systems.
Technical details
-
Hybrid Retrieval
1052s
A combination of semantic search and keyword-based search (like BM25) is recommended for enterprise use cases, providing better results than either method alone. Re-ranking the top results using an LLM further improves accuracy.
-
File Traversal Primitives
1157s
Agents should be equipped with tools like file listing, metadata filtering, and `grep` commands to allow them to narrow down the search scope and ground answers without needing to download the entire corpus.
-
Parsing and Indexing Pipeline
The recommended pipeline involves three stages: 1) Ingestion (getting documents into a system); 2) Parsing (converting documents into a standardized format, e.g., Markdown); and 3) Re-indexing (making the data available for retrieval).
-
Storage Architecture
For large-scale, multi-tenant systems, it is recommended to persist data onto disk and use a shielded memory layer for efficient retrieval, rather than relying solely on in-memory vector stores.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.