# Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex

## Executive summary

This talk details the evolution of document search, contrasting the simple, text-based nature of codebases (where plain file traversal and `grep` suffice) with the complexity of enterprise data (multimodal PDFs, schematics, and slides). The solution presented is an agentic 'harness' that allows an AI agent to dynamically choose the best retrieval method—from hybrid search to metadata filtering and specialized document reading—to ground answers from massive, messy, and permissioned corporate data sets. Key components include hybrid retrieval, structured parsing via LlamaParse/LiteParse, and managing production concerns like multi-tenancy and data freshness.

## Key takeaways

- Code vs. Company Data Search: Plain file traversal (like Claude Code's approach) works well for code because it is small, text-based, and structured. However, company data (PDFs, schematics, etc.) is huge, multimodal, and lacks a simple folder hierarchy, necessitating advanced retrieval methods.
- The Agentic Search Harness: Instead of relying on a single method, agents should be given a suite of tools (a 'harness') to choose from: hybrid search (semantic + keyword), metadata filtering, file listing/traversal, and specialized document reading.
- Handling Complex Documents: For multimodal documents, specialized parsing is required. Tools like LlamaParse and LiteParse are recommended to extract structured content (tables, charts) and generate page screenshots, which aids in spatial reasoning and improves Retrieval-Augmented Generation (RAG).
- Production Scaling Concerns: When deploying in production, engineers must address multi-tenancy, granular permissioning (e.g., pulling metadata from Google Drive/SharePoint), and data freshness/synchronization, which are major challenges for pre-indexing systems.

## Technical details

- Hybrid Retrieval: A combination of semantic search and keyword-based search (like BM25) is recommended for enterprise use cases, providing better results than either method alone. Re-ranking the top results using an LLM further improves accuracy.
- File Traversal Primitives: Agents should be equipped with tools like file listing, metadata filtering, and `grep` commands to allow them to narrow down the search scope and ground answers without needing to download the entire corpus.
- Parsing and Indexing Pipeline: The recommended pipeline involves three stages: 1) Ingestion (getting documents into a system); 2) Parsing (converting documents into a standardized format, e.g., Markdown); and 3) Re-indexing (making the data available for retrieval).
- Storage Architecture: For large-scale, multi-tenant systems, it is recommended to persist data onto disk and use a shielded memory layer for efficient retrieval, rather than relying solely on in-memory vector stores.

## Practical implications

- Build engineers must design retrieval systems that are modular, allowing the agent to select the optimal search tool (e.g., metadata filter vs. semantic search) based on the query and data type.
- Scaling requires moving beyond simple local file systems and implementing robust, multi-stage pipelines (Ingest $\rightarrow$ Parse $\rightarrow$ Index) that handle complex data types and maintain data freshness.
- The system must incorporate permissioning and multi-tenancy metadata at the vector storage layer to ensure data isolation and security.

## Topics

Agentic Orchestration, RAG, Document Parsing, Hybrid Search, Knowledge Management, Scalability, LiteParse, LlamaParse, LlamaIndex

Source: https://www.youtube.com/watch?v=X4w2Pkz5tDY
