AI Engineer

Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust

Published 2026-10-05 · Duration 18:15

Summary

This talk details a real-world evaluation comparing agentic search against vector search for coding agents. Using merged fix PRs from Microsoft's TypeScript Go repo, the speaker found that both methods achieved the same accuracy in locating buggy code. However, vector search was significantly more expensive (four times the cost) because its chunk-based approach lacked the surrounding 'connective tissue' and context necessary for the agent to solve the bug efficiently, unlike agentic search, which mimics human code exploration using tools like `grep` and `find`.

Download summary

Key takeaways

  1. Evals are a Team Sport 12:27

    Creating robust evaluations requires collaboration between AI engineers, Product Managers (for developing hypotheses), Subject Matter Experts (for labeling ground truth), and Data Analysts.

  2. Vector Search vs. Agentic Search

    Vector search returns semantically similar code chunks but often misses critical context (like imports or calling code). Agentic search, by using tools (e.g., `grep`, `find`), can follow the connective logic across files, mimicking human debugging.

  3. Cost Efficiency is Critical

    The evaluation demonstrated that while vector search achieved high accuracy, its constant need for multiple searches made it four times more expensive than agentic search.

Technical details

  • Evaluation Methodology 302s

    Evals require four components: a Dataset (including golden standard, edge cases, and failure modes), a Task (defining system prompts and the model), a Scoring System (e.g., deterministic scoring, LLM as a judge), and an Experiment (a specific configuration of the three above).

  • Agentic Search Implementation 1042s

    Agentic search allows an LLM to explore a codebase like a human, using tools such as `grep`, `find`, `ls`, and `cat` to follow function calls and read files sequentially.

  • Vector Search Implementation 932s

    Vector search converts code/text into embeddings (e.g., using a vector database like Qdrant or Pinecone). It returns code chunks that semantically match the query, but these chunks may lack surrounding context.

  • Evaluation Setup

    The evaluation used merged 'fix' PRs from Microsoft's TypeScript Go repo. The task involved having Claude Code locate the buggy code based on the diff between the buggy and fixed versions. The scoring was binary: 100% if the test suite passed, 0% if it failed.

Mentioned resources

  • Braintrust (Company)
  • Microsoft's TypeScript Go repo (Codebase)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.