# DGX Spark Live: Smart Routing for Hybrid AI

## Executive summary

This presentation details a solution for managing the escalating costs of AI in enterprise environments, particularly due to the complexity of Agentic AI. The core concept is 'smart routing' for hybrid AI, which intelligently directs workloads between local, on-premise resources (like the Lenovo ThinkStation PGX) and the cloud. By processing tasks locally, organizations can significantly reduce reliance on expensive cloud API calls, thereby regaining control over their IT AI budget and maximizing local compute capacity.

## Key takeaways

- Addressing AI Tokenomics Costs: Agentic AI significantly increases token usage (e.g., from a few hundred tokens to tens of thousands) because the context window must be wrapped with instructions and tools. Since cloud models are charged per million tokens, this rapid increase in usage can quickly deplete annual AI budgets.
- The Hybrid AI Architecture: The system uses the open-source Nvidia Switchyard Router, guided by a small Judge Model (0.8B parameters). This router determines the optimal destination for a query—local efficient model, local strong model, or the cloud—based on complexity and required resources.
- Significant Cost Savings: By routing tasks locally, the system demonstrated substantial cost savings. The speaker calculated that a local run saved an estimated $4,500 in potential cloud token costs, demonstrating a rapid payback period for the local hardware.

## Technical details

- Hardware Platform: The demonstration uses a Lenovo ThinkStation PGX, powered by DGX Spark and based on the NVIDIA GB10, functioning as a local 'AI data center in a backpack.'
- Model Architecture and Efficiency: Three models run concurrently: 1) Neatron 3.5 Lightning (efficient, 3B active parameters out of 30B total) for general tasks; 2) Qwen 3.8B (strong, dense, multimodal) for image recognition and complex tasks; 3) Cloud endpoint for highly complex, multi-stage queries.
- Smart Routing Mechanism: The Nvidia Switchyard Router and Judge Model manage the workflow. The Judge Model assesses the context and routes the request. Escalation occurs if a model hallucinates or exceeds its context window, moving the task to the next, more powerful model (e.g., Efficient $ ightarrow$ Strong $ ightarrow$ Cloud).
- Data Security: The cloud routing path includes a redaction portion (markdown file control) to remove sensitive data like HIPAA, GDPR, social security numbers, and credit card numbers before transmission.

## Practical implications

- Organizations can significantly reduce operational IT expenditure by processing routine AI queries locally, reserving expensive cloud tokens for truly complex tasks.
- The system allows for the segregation of AI resources, dedicating local capacity to core business functions while maintaining user access to AI.
- The PGX unit can serve as a versatile 'AI companion device' or a dedicated local inference server, enabling productivity on the road or in a data center.
- The architecture provides a clear path for capacity planning by quantifying the cost savings achieved through local processing.

## Topics

AI Cost Management, Edge AI, Hybrid Cloud Computing, LLM Routing, Workstation Computing, Lenovo ThinkStation PGX, DGX Spark, NVIDIA Switchyard Router, NVFP4 quantization, Neatron 3.5 Lightning, Qwen 3.8B, NVIDIA GB10

Source: https://www.youtube.com/watch?v=XY6ySMd2Lzs
