AI Engineer

Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex

Published 2026-10-07 · Duration 23:22

Summary

This talk details the evolution of document search, contrasting the simple, text-based nature of codebases (where plain file traversal and `grep` suffice) with the complexity of enterprise data (multimodal PDFs, schematics, and slides). The solution presented is an agentic 'harness' that allows an AI agent to dynamically choose the best retrieval method—from hybrid search to metadata filtering and specialized document reading—to ground answers from massive, messy, and permissioned corporate data sets. Key components include hybrid retrieval, structured parsing via LlamaParse/LiteParse, and managing production concerns like multi-tenancy and data freshness.

Download summary

Key takeaways

  1. Code vs. Company Data Search 5:27

    Plain file traversal (like Claude Code's approach) works well for code because it is small, text-based, and structured. However, company data (PDFs, schematics, etc.) is huge, multimodal, and lacks a simple folder hierarchy, necessitating advanced retrieval methods.

  2. The Agentic Search Harness 13:42

    Instead of relying on a single method, agents should be given a suite of tools (a 'harness') to choose from: hybrid search (semantic + keyword), metadata filtering, file listing/traversal, and specialized document reading.

  3. Handling Complex Documents 22:07

    For multimodal documents, specialized parsing is required. Tools like LlamaParse and LiteParse are recommended to extract structured content (tables, charts) and generate page screenshots, which aids in spatial reasoning and improves Retrieval-Augmented Generation (RAG).

  4. Production Scaling Concerns 20:42

    When deploying in production, engineers must address multi-tenancy, granular permissioning (e.g., pulling metadata from Google Drive/SharePoint), and data freshness/synchronization, which are major challenges for pre-indexing systems.

Technical details

  • Hybrid Retrieval 1052s

    A combination of semantic search and keyword-based search (like BM25) is recommended for enterprise use cases, providing better results than either method alone. Re-ranking the top results using an LLM further improves accuracy.

  • File Traversal Primitives 1157s

    Agents should be equipped with tools like file listing, metadata filtering, and `grep` commands to allow them to narrow down the search scope and ground answers without needing to download the entire corpus.

  • Parsing and Indexing Pipeline

    The recommended pipeline involves three stages: 1) Ingestion (getting documents into a system); 2) Parsing (converting documents into a standardized format, e.g., Markdown); and 3) Re-indexing (making the data available for retrieval).

  • Storage Architecture

    For large-scale, multi-tenant systems, it is recommended to persist data onto disk and use a shielded memory layer for efficient retrieval, rather than relying solely on in-memory vector stores.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.