LangChain

How to Build a Model Router in the Harness

Published 2026-10-06 · Duration 13:39

Summary

This talk outlines the methodology for building a model router within an agent's harness to significantly reduce operational costs without sacrificing quality. The process involves four key steps: analyzing agent task usage via LangSmith traces, selecting models along the Pareto frontier, implementing the router (e.g., using Jev), and rigorously tracking outcomes via A/B testing. In an experiment using Open SWE, the implemented router achieved a 64% reduction in median cost per thread.

Download summary

Key takeaways

  1. Cost Reduction via Model Routing 3:30

    In an A/B test against always using the performance model, the model router cut median cost per thread by 64% with no measurable change in quality, demonstrating that cheaper models can handle a large proportion of tasks.

  2. Model Selection Strategy 7:33

    Models should be selected along the Pareto frontier—the set of optimal models balancing benchmark accuracy, evals, cost, and latency. The strategy involves assigning tasks to appropriate tiers (fast, balanced, performance).

  3. Router Implementation 10:49

    The router can be powered by specialized decision models, such as Jev, which were found to be significantly faster and cheaper than traditional LLMs for constrained decision-making.

Technical details

  • Step 1: Task Mapping (LangSmith) 240s

    Analyze agent traces (via LangSmith) to understand the distribution of tasks (e.g., code changes, bug fixes, design questions). This analysis determines the task space and informs the criteria for the router.

  • Step 2: Model Selection (Pareto Frontier) 453s

    Identify models that offer optimal trade-offs between cost and intelligence. The process involves selecting models for different tiers (e.g., GLM-5.3-Flash for fast, GPT-5.6 Sol for balanced, GPT-6 Astra for performance).

  • Step 3: Building the Router in the Harness 559s

    The router is integrated into the agent's harness, typically running once at the start of a thread to select the appropriate model. The router's criteria are derived from task analysis and official provider guides.

  • Step 4: Outcome Tracking 746s

    Measure success using either offline evaluations (safe but difficult to set up) or A/B testing on live traffic. Key metrics tracked include the merged PR rate and user feedback.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.