# Ask the Experts: How NeMo Switchyard Helps Agents Select Models | Nemotron Labs

## Executive summary

NeMo Switchyard is an open-source model routing library designed for AI agents to solve the problem of relying on a single monolithic LLM. It automatically routes each agent query or step to the optimal model—selecting from any combination of local/cloud and open/closed models—based on real-time needs, optimizing for accuracy, cost, and latency. The system operates beyond simple request routing by tracking state across multi-turn agentic workflows, making it a critical component for building robust, efficient AI systems.

## Key takeaways

- System of Models Approach: The industry is moving away from the 'one model to rule them all' concept toward a 'system of models,' where multiple specialized models are used for different tasks, improving efficiency and capability (0:02:45).
- Agent-Aware Routing vs. Simple Routing: Switchyard is more than a simple router; it operates on an agentic workflow, tracking state (e.g., tool calls, message history) across multi-turn sessions to make intelligent model selection decisions (0:04:25).
- Optimization and Learning: The system treats model selection as an optimization problem. It can learn by analyzing agent traces and behavior, predicting potential errors or resource needs to route proactively and save tokens/time (0:21:58).
- Full-Stack Routing Flywheel: The roadmap envisions a full 'flywheel' of routing, connecting model selection to inference optimization (via NVIDIA Dynamo) and data privacy/anonymization. This allows for continuous improvement across the entire agent lifecycle (0:06:10).

## Technical details

- Architecture & Integration: Switchyard can be integrated in multiple ways: directly into the agent harness/application, embedded natively, or within an existing LLM gateway (like OpenRouter). It operates on a provider-neutral format for algorithm development (0:03:58; 0:04:29).
- Optimization Parameters: Users can tune routing strategies based on specific priorities, including accuracy, cost (token budget), and latency. Using routing algorithms has been shown to yield significant token cost reductions (50-80%) while maintaining frontier model accuracy (0:13:49).
- Observability and Telemetry: Switchyard supports OpenTelemetry (OTEL) for capturing performance metrics. Future development aims to expose not just *what* model was chosen, but *why* the decision was made by the routing algorithm (0:16:27).
- Advanced Routing Algorithms: The library includes specialized algorithms like 'escalation router' (for handling persistent errors) and 'stage router' (which uses tool call history to inform model selection), allowing for complex, multi-signal decision making (0:23:15; 0:24:39).

## Practical implications

- Build engineers can design highly efficient agent pipelines by abstracting model selection logic into a dedicated routing layer (Switchyard).
- The ability to optimize for cost and latency allows teams to deploy powerful agents using local, low-cost models without sacrificing the performance of frontier cloud models.
- Integrating Switchyard provides a standardized way to manage complex multi-model workflows, simplifying agent development and maintenance.

## Topics

AI Agents, Model Routing, LLM Optimization, System Architecture, Open Source AI Tools, NeMo Switchyard GitHub, NVIDIA Dynamo

Source: https://www.youtube.com/watch?v=tSPAckZON0A
