# AI Tokenomics Explained

## Executive summary

The economics of AI inference are rapidly changing due to the rise of agentic AI, which significantly increases token usage compared to basic chat. To scale AI profitably, organizations must implement an extreme co-design approach across the entire stack: model, system, and software. Key strategies include optimizing token utility by matching model capability to the specific use case, accurately forecasting token demand using workload multipliers, and maximizing token supply through advanced architectures like Mixture of Experts (MoE) and specialized interconnects (e.g., NVL72).

## Key takeaways

- Agentic AI dramatically increases token usage.: Agentic workflows, which involve multiple turns between AI, sub-agents, and tool calls (e.g., querying databases, running Python scripts), can use up to 15x more tokens than basic chat interactions. This makes the economics of AI inference critical.
- Token value is determined by intelligence, context, and interactivity.: The value of a token depends on the intelligence carried by the model, the length of the context it can process, and the speed of generation (interactivity). Using premium tokens for simple use cases leads to poor economics.
- Optimization requires co-design across Model, System, and Software.: Achieving optimal token cost and maximum output requires integrating model efficiency (e.g., Mixture of Experts), system efficiency (e.g., NVLink/NVL72 for fast inter-GPU communication), and software efficiency (e.g., Dynamo for KV cache offloading and routing).
- Forecasting token demand requires workload multipliers.: Beyond basic user count and requests, demand forecasting must account for workload multipliers, such as reasoning tokens (thinking tokens not seen by the user), agentic loops, and high cache hit rates (e.g., in enterprise search).

## Technical details

- Agentic AI Workflows: Agentic AI involves complex, multi-step processes where a main agent coordinates tool calls (running on CPU) and sub-agents. This multi-turn interaction is the primary driver of increased token consumption.
- Mixture of Experts (MoE): MoE is an efficient model architecture where only relevant parameters are activated for a given token, allowing the model to achieve high intelligence while lowering per-token compute and HBM bandwidth requirements. However, realizing its full benefit requires fast inter-GPU communication.
- System Efficiency and Interconnects: To support MoE, system efficiency is vital. Technologies like the Nvidia NVLink and NVL72 system connect 72 GPUs, providing ultra-fast scale-up interconnects that outperform Ethernet by 3x in latency and 10x in packet rate. This system also supports in-network compute for combining expert results.
- Software Optimization Stack: Software efficiency enables critical optimizations like quantization, KV cache offloading, KV-aware routing, and prefix caching. The Dynamo serving and orchestration layer is key for managing these complex, multi-model workflows.
- Workload Benchmarking: Capacity planning should use a multi-layered approach: 1) Base demand (Users * Requests * Tokens/Request); 2) Workload multipliers (reasoning tokens, agentic loops); and 3) Operational factors (seasonal variation, growth rate).

## Practical implications

- When prototyping, use frontier capability to validate use cases, but plan for production deployment using cost-optimized, right-sized models.
- Implement real-time request routing (e.g., using Nemo Switchyard) to direct requests to the most cost-effective model based on policy, pricing, and latency requirements.
- Establish robust governance and observability (control plane) to track token usage and identify areas for optimization.
- For capacity planning, model peak demand and burst capacity rather than relying solely on average daily usage.

## Topics

AI Inference Economics, Large Language Models (LLMs), AI Infrastructure, Tokenomics, System Architecture, NVIDIA Agent Toolkit, Nemo Switchyard, Neotron 3 Family, AT&T tokconomics paper

Source: https://www.youtube.com/watch?v=VlLd3PsQ7fc
