# Every AI Company Is Accidentally Building a Bank — Dor Sasson, Stigg

## Executive summary

The talk argues that modern AI product scaling requires treating AI consumption not merely as a usage metric, but as a complex financial system. Current architectures often fail because they check entitlements and usage *after* the inference (post-invoice), leading to massive overspends and operational crises (e.g., Anthropic, OpenAI pricing changes). The solution involves implementing financial infrastructure that performs synchronous checks *before* consumption, similar to an ATM withdrawal, and settling asynchronously.

## Key takeaways

- AI Pricing Emergencies are Infrastructure Failures: Recent pricing scrambles (Anthropic, OpenAI, GitHub) are not just commercial issues; they expose fundamental architectural flaws where usage checks happen after the invoice, making scaling impossible.
- Synchronous Check, Asynchronous Settle: The core architectural shift needed is to enforce entitlements and check balances synchronously at the request level (the 'hot path'), while reconciling and settling the costs asynchronously.
- Adopting Banking Principles: Scaling AI requires implementing financial concepts like double-entry bookkeeping, concurrency control, hold-and-settle mechanisms, and managing multiple, distinct credit pools (e.g., debit, cash, savings).
- Visibility is Table Stakes: Enterprise clients now demand fine-grained visibility into AI workload consumption across different models, teams, and agents, making this capability a non-negotiable requirement for doing business.

## Technical details

- Billing Architecture Flaw: Most current systems check and settle usage *after* inference. This 'after-effect' logging leads to overspends and poor business practices. The required architecture must perform decision-making (checking entitlements, rate limits) immediately at runtime.
- OpenAI's Financial Engineering: OpenAI's internal architecture requires a decision waterfall that calculates user/agent entitlements and rate limits immediately at the front end, confirming if consumption is allowed *before* the draw down can occur.
- Concurrency and Double Spend: When multiple agents or parties access a shared resource pool, concurrency becomes a critical issue. Financial systems solve this using mechanisms like double-entry bookkeeping to prevent simultaneous over-draws.
- Hierarchical Spend Management: Credit pools are no longer flat. Enterprise sales require managing budgets and spend caps across complex hierarchies (C-suite, teams, users, agents), demanding sophisticated governance and allocation logic.

## Practical implications

- AI companies must redesign their billing and usage tracking layers to function like financial institutions, moving from simple usage logging to complex, real-time credit management.
- The shift requires building dedicated financial infrastructure capable of handling synchronous checks, concurrency, and multi-source credit allocation.
- Focusing on 'reserve-and-settle' patterns is crucial for selling AI workloads at enterprise scale.

## Topics

AI Infrastructure, Financial Engineering, Billing Systems, Distributed Systems, Scaling, Stigg, Anthropic, OpenAI, GitHub Copilot, OpenClaw

Source: https://www.youtube.com/watch?v=cf2IhzqeQH4
