# Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust

## Executive summary

This talk details a real-world evaluation comparing agentic search against vector search for coding agents. Using merged fix PRs from Microsoft's TypeScript Go repo, the speaker found that both methods achieved the same accuracy in locating buggy code. However, vector search was significantly more expensive (four times the cost) because its chunk-based approach lacked the surrounding 'connective tissue' and context necessary for the agent to solve the bug efficiently, unlike agentic search, which mimics human code exploration using tools like `grep` and `find`.

## Key takeaways

- Evals are a Team Sport: Creating robust evaluations requires collaboration between AI engineers, Product Managers (for developing hypotheses), Subject Matter Experts (for labeling ground truth), and Data Analysts.
- Vector Search vs. Agentic Search: Vector search returns semantically similar code chunks but often misses critical context (like imports or calling code). Agentic search, by using tools (e.g., `grep`, `find`), can follow the connective logic across files, mimicking human debugging.
- Cost Efficiency is Critical: The evaluation demonstrated that while vector search achieved high accuracy, its constant need for multiple searches made it four times more expensive than agentic search.

## Technical details

- Evaluation Methodology: Evals require four components: a Dataset (including golden standard, edge cases, and failure modes), a Task (defining system prompts and the model), a Scoring System (e.g., deterministic scoring, LLM as a judge), and an Experiment (a specific configuration of the three above).
- Agentic Search Implementation: Agentic search allows an LLM to explore a codebase like a human, using tools such as `grep`, `find`, `ls`, and `cat` to follow function calls and read files sequentially.
- Vector Search Implementation: Vector search converts code/text into embeddings (e.g., using a vector database like Qdrant or Pinecone). It returns code chunks that semantically match the query, but these chunks may lack surrounding context.
- Evaluation Setup: The evaluation used merged 'fix' PRs from Microsoft's TypeScript Go repo. The task involved having Claude Code locate the buggy code based on the diff between the buggy and fixed versions. The scoring was binary: 100% if the test suite passed, 0% if it failed.

## Practical implications

- When designing AI systems, prioritize evaluating cost efficiency alongside accuracy, as complex search methods can lead to excessive API calls.
- For code-related tasks, consider agentic approaches that allow for multi-step, tool-based reasoning over simple semantic chunk retrieval.
- Implement robust observability (e.g., passing parent span IDs) to ensure full visibility into subprocesses and complex agent runs for effective debugging.
- Use evals as a continuous feedback loop: Production logs -> Dataset -> Evals -> Code Change -> Production Logs.

## Topics

AI Evaluation, Code Generation, LLM Agents, Vector Databases, Software Testing, Observability, Braintrust, Microsoft's TypeScript Go repo

Source: https://www.youtube.com/watch?v=T3SS931wU0I
