Every AI Company Is Accidentally Building a Bank — Dor Sasson, Stigg
The talk argues that modern AI product scaling requires treating AI consumption not merely as a usage metric, but as a complex financial system. Current architectures often fail because they check entitlements and usage *after* the inference (post-invoice), leading to massive overspends and operational crises (e.g., Anthropic, OpenAI pricing changes). The solution involves implementing financial infrastructure that performs synchronous checks *before* consumption, similar to an ATM withdrawal, and settling asynchronously.
Key takeaways
-
AI Pricing Emergencies are Infrastructure Failures
2:00
Recent pricing scrambles (Anthropic, OpenAI, GitHub) are not just commercial issues; they expose fundamental architectural flaws where usage checks happen after the invoice, making scaling impossible.
-
Synchronous Check, Asynchronous Settle
10:27
The core architectural shift needed is to enforce entitlements and check balances synchronously at the request level (the 'hot path'), while reconciling and settling the costs asynchronously.
-
Adopting Banking Principles
17:11
Scaling AI requires implementing financial concepts like double-entry bookkeeping, concurrency control, hold-and-settle mechanisms, and managing multiple, distinct credit pools (e.g., debit, cash, savings).
-
Visibility is Table Stakes
Enterprise clients now demand fine-grained visibility into AI workload consumption across different models, teams, and agents, making this capability a non-negotiable requirement for doing business.