How to Build a Model Router in the Harness
This talk outlines the methodology for building a model router within an agent's harness to significantly reduce operational costs without sacrificing quality. The process involves four key steps: analyzing agent task usage via LangSmith traces, selecting models along the Pareto frontier, implementing the router (e.g., using Jev), and rigorously tracking outcomes via A/B testing. In an experiment using Open SWE, the implemented router achieved a 64% reduction in median cost per thread.
Key takeaways
-
Cost Reduction via Model Routing
3:30
In an A/B test against always using the performance model, the model router cut median cost per thread by 64% with no measurable change in quality, demonstrating that cheaper models can handle a large proportion of tasks.
-
Model Selection Strategy
7:33
Models should be selected along the Pareto frontier—the set of optimal models balancing benchmark accuracy, evals, cost, and latency. The strategy involves assigning tasks to appropriate tiers (fast, balanced, performance).
-
Router Implementation
10:49
The router can be powered by specialized decision models, such as Jev, which were found to be significantly faster and cheaper than traditional LLMs for constrained decision-making.