How to Build a Model Router in the Harness
Summary
This talk outlines the methodology for building a model router within an agent's harness to significantly reduce operational costs without sacrificing quality. The process involves four key steps: analyzing agent task usage via LangSmith traces, selecting models along the Pareto frontier, implementing the router (e.g., using Jev), and rigorously tracking outcomes via A/B testing. In an experiment using Open SWE, the implemented router achieved a 64% reduction in median cost per thread.
Key takeaways
-
Cost Reduction via Model Routing
3:30
In an A/B test against always using the performance model, the model router cut median cost per thread by 64% with no measurable change in quality, demonstrating that cheaper models can handle a large proportion of tasks.
-
Model Selection Strategy
7:33
Models should be selected along the Pareto frontier—the set of optimal models balancing benchmark accuracy, evals, cost, and latency. The strategy involves assigning tasks to appropriate tiers (fast, balanced, performance).
-
Router Implementation
10:49
The router can be powered by specialized decision models, such as Jev, which were found to be significantly faster and cheaper than traditional LLMs for constrained decision-making.
Technical details
-
Step 1: Task Mapping (LangSmith)
240s
Analyze agent traces (via LangSmith) to understand the distribution of tasks (e.g., code changes, bug fixes, design questions). This analysis determines the task space and informs the criteria for the router.
-
Step 2: Model Selection (Pareto Frontier)
453s
Identify models that offer optimal trade-offs between cost and intelligence. The process involves selecting models for different tiers (e.g., GLM-5.3-Flash for fast, GPT-5.6 Sol for balanced, GPT-6 Astra for performance).
-
Step 3: Building the Router in the Harness
559s
The router is integrated into the agent's harness, typically running once at the start of a thread to select the appropriate model. The router's criteria are derived from task analysis and official provider guides.
-
Step 4: Outcome Tracking
746s
Measure success using either offline evaluations (safe but difficult to set up) or A/B testing on live traffic. Key metrics tracked include the merged PR rate and user feedback.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.