Topic

Hybrid Search

All digests tagged Hybrid Search

Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex thumbnail

· 23:22

Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex

This talk details the evolution of document search, contrasting the simple, text-based nature of codebases (where plain file traversal and `grep` suffice) with the complexity of enterprise data (multimodal PDFs, schematics, and slides). The solution presented is an agentic 'harness' that allows an AI agent to dynamically choose the best retrieval method—from hybrid search to metadata filtering and specialized document reading—to ground answers from massive, messy, and permissioned corporate data sets. Key components include hybrid retrieval, structured parsing via LlamaParse/LiteParse, and managing production concerns like multi-tenancy and data freshness.

Key takeaways

  1. Code vs. Company Data Search 5:27

    Plain file traversal (like Claude Code's approach) works well for code because it is small, text-based, and structured. However, company data (PDFs, schematics, etc.) is huge, multimodal, and lacks a simple folder hierarchy, necessitating advanced retrieval methods.

  2. The Agentic Search Harness 13:42

    Instead of relying on a single method, agents should be given a suite of tools (a 'harness') to choose from: hybrid search (semantic + keyword), metadata filtering, file listing/traversal, and specialized document reading.

  3. Handling Complex Documents 22:07

    For multimodal documents, specialized parsing is required. Tools like LlamaParse and LiteParse are recommended to extract structured content (tables, charts) and generate page screenshots, which aids in spatial reasoning and improves Retrieval-Augmented Generation (RAG).

  4. Production Scaling Concerns 20:42

    When deploying in production, engineers must address multi-tenancy, granular permissioning (e.g., pulling metadata from Google Drive/SharePoint), and data freshness/synchronization, which are major challenges for pre-indexing systems.

Watch on YouTube Full article