# How to Build a Model Router in the Harness

## Executive summary

This talk outlines the methodology for building a model router within an agent's harness to significantly reduce operational costs without sacrificing quality. The process involves four key steps: analyzing agent task usage via LangSmith traces, selecting models along the Pareto frontier, implementing the router (e.g., using Jev), and rigorously tracking outcomes via A/B testing. In an experiment using Open SWE, the implemented router achieved a 64% reduction in median cost per thread.

## Key takeaways

- Cost Reduction via Model Routing: In an A/B test against always using the performance model, the model router cut median cost per thread by 64% with no measurable change in quality, demonstrating that cheaper models can handle a large proportion of tasks.
- Model Selection Strategy: Models should be selected along the Pareto frontier—the set of optimal models balancing benchmark accuracy, evals, cost, and latency. The strategy involves assigning tasks to appropriate tiers (fast, balanced, performance).
- Router Implementation: The router can be powered by specialized decision models, such as Jev, which were found to be significantly faster and cheaper than traditional LLMs for constrained decision-making.

## Technical details

- Step 1: Task Mapping (LangSmith): Analyze agent traces (via LangSmith) to understand the distribution of tasks (e.g., code changes, bug fixes, design questions). This analysis determines the task space and informs the criteria for the router.
- Step 2: Model Selection (Pareto Frontier): Identify models that offer optimal trade-offs between cost and intelligence. The process involves selecting models for different tiers (e.g., GLM-5.3-Flash for fast, GPT-5.6 Sol for balanced, GPT-6 Astra for performance).
- Step 3: Building the Router in the Harness: The router is integrated into the agent's harness, typically running once at the start of a thread to select the appropriate model. The router's criteria are derived from task analysis and official provider guides.
- Step 4: Outcome Tracking: Measure success using either offline evaluations (safe but difficult to set up) or A/B testing on live traffic. Key metrics tracked include the merged PR rate and user feedback.

## Practical implications

- Implement model routing middleware to dynamically select the most cost-effective model for each task within an agent's workflow.
- Use tracing tools (like LangSmith) to perform exploratory data analysis on agent task types and complexity before building the router.
- When designing agent systems, treat model selection as a context engineering problem and integrate the router directly into the agent's core harness.
- Measure the impact of routing decisions using A/B testing on live user traffic to ensure cost savings do not degrade core performance metrics (e.g., PR merge rate).

## Topics

Agent Orchestration, LLM Cost Optimization, Model Routing, A/B Testing, LangChain, Companion blog, How to Build a Model Router in the Harness, Building a Harness with Jev, Open SWE, LangSmith

Source: https://www.youtube.com/watch?v=5kTFyEOgark
