AI Tokenomics Explained
Summary
The economics of AI inference are rapidly changing due to the rise of agentic AI, which significantly increases token usage compared to basic chat. To scale AI profitably, organizations must implement an extreme co-design approach across the entire stack: model, system, and software. Key strategies include optimizing token utility by matching model capability to the specific use case, accurately forecasting token demand using workload multipliers, and maximizing token supply through advanced architectures like Mixture of Experts (MoE) and specialized interconnects (e.g., NVL72).
Key takeaways
-
Agentic AI dramatically increases token usage.
2:00
Agentic workflows, which involve multiple turns between AI, sub-agents, and tool calls (e.g., querying databases, running Python scripts), can use up to 15x more tokens than basic chat interactions. This makes the economics of AI inference critical.
-
Token value is determined by intelligence, context, and interactivity.
4:10
The value of a token depends on the intelligence carried by the model, the length of the context it can process, and the speed of generation (interactivity). Using premium tokens for simple use cases leads to poor economics.
-
Optimization requires co-design across Model, System, and Software.
12:30
Achieving optimal token cost and maximum output requires integrating model efficiency (e.g., Mixture of Experts), system efficiency (e.g., NVLink/NVL72 for fast inter-GPU communication), and software efficiency (e.g., Dynamo for KV cache offloading and routing).
-
Forecasting token demand requires workload multipliers.
16:40
Beyond basic user count and requests, demand forecasting must account for workload multipliers, such as reasoning tokens (thinking tokens not seen by the user), agentic loops, and high cache hit rates (e.g., in enterprise search).
Technical details
-
Agentic AI Workflows
120s
Agentic AI involves complex, multi-step processes where a main agent coordinates tool calls (running on CPU) and sub-agents. This multi-turn interaction is the primary driver of increased token consumption.
-
Mixture of Experts (MoE)
650s
MoE is an efficient model architecture where only relevant parameters are activated for a given token, allowing the model to achieve high intelligence while lowering per-token compute and HBM bandwidth requirements. However, realizing its full benefit requires fast inter-GPU communication.
-
System Efficiency and Interconnects
700s
To support MoE, system efficiency is vital. Technologies like the Nvidia NVLink and NVL72 system connect 72 GPUs, providing ultra-fast scale-up interconnects that outperform Ethernet by 3x in latency and 10x in packet rate. This system also supports in-network compute for combining expert results.
-
Software Optimization Stack
850s
Software efficiency enables critical optimizations like quantization, KV cache offloading, KV-aware routing, and prefix caching. The Dynamo serving and orchestration layer is key for managing these complex, multi-model workflows.
-
Workload Benchmarking
1000s
Capacity planning should use a multi-layered approach: 1) Base demand (Users * Requests * Tokens/Request); 2) Workload multipliers (reasoning tokens, agentic loops); and 3) Operational factors (seasonal variation, growth rate).
Mentioned resources
- NVIDIA Agent Toolkit
- Nemo Switchyard
- Neotron 3 Family
- AT&T tokconomics paper
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.