Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust
This talk details a real-world evaluation comparing agentic search against vector search for coding agents. Using merged fix PRs from Microsoft's TypeScript Go repo, the speaker found that both methods achieved the same accuracy in locating buggy code. However, vector search was significantly more expensive (four times the cost) because its chunk-based approach lacked the surrounding 'connective tissue' and context necessary for the agent to solve the bug efficiently, unlike agentic search, which mimics human code exploration using tools like `grep` and `find`.
Key takeaways
-
Evals are a Team Sport
12:27
Creating robust evaluations requires collaboration between AI engineers, Product Managers (for developing hypotheses), Subject Matter Experts (for labeling ground truth), and Data Analysts.
-
Vector Search vs. Agentic Search
Vector search returns semantically similar code chunks but often misses critical context (like imports or calling code). Agentic search, by using tools (e.g., `grep`, `find`), can follow the connective logic across files, mimicking human debugging.
-
Cost Efficiency is Critical
The evaluation demonstrated that while vector search achieved high accuracy, its constant need for multiple searches made it four times more expensive than agentic search.