AI Engineer

Every AI Company Is Accidentally Building a Bank — Dor Sasson, Stigg

Published 2026-10-07 · Duration 20:08

Summary

The talk argues that modern AI product scaling requires treating AI consumption not merely as a usage metric, but as a complex financial system. Current architectures often fail because they check entitlements and usage *after* the inference (post-invoice), leading to massive overspends and operational crises (e.g., Anthropic, OpenAI pricing changes). The solution involves implementing financial infrastructure that performs synchronous checks *before* consumption, similar to an ATM withdrawal, and settling asynchronously.

Download summary

Key takeaways

  1. AI Pricing Emergencies are Infrastructure Failures 2:00

    Recent pricing scrambles (Anthropic, OpenAI, GitHub) are not just commercial issues; they expose fundamental architectural flaws where usage checks happen after the invoice, making scaling impossible.

  2. Synchronous Check, Asynchronous Settle 10:27

    The core architectural shift needed is to enforce entitlements and check balances synchronously at the request level (the 'hot path'), while reconciling and settling the costs asynchronously.

  3. Adopting Banking Principles 17:11

    Scaling AI requires implementing financial concepts like double-entry bookkeeping, concurrency control, hold-and-settle mechanisms, and managing multiple, distinct credit pools (e.g., debit, cash, savings).

  4. Visibility is Table Stakes

    Enterprise clients now demand fine-grained visibility into AI workload consumption across different models, teams, and agents, making this capability a non-negotiable requirement for doing business.

Technical details

  • Billing Architecture Flaw 502s

    Most current systems check and settle usage *after* inference. This 'after-effect' logging leads to overspends and poor business practices. The required architecture must perform decision-making (checking entitlements, rate limits) immediately at runtime.

  • OpenAI's Financial Engineering 811s

    OpenAI's internal architecture requires a decision waterfall that calculates user/agent entitlements and rate limits immediately at the front end, confirming if consumption is allowed *before* the draw down can occur.

  • Concurrency and Double Spend 1031s

    When multiple agents or parties access a shared resource pool, concurrency becomes a critical issue. Financial systems solve this using mechanisms like double-entry bookkeeping to prevent simultaneous over-draws.

  • Hierarchical Spend Management

    Credit pools are no longer flat. Enterprise sales require managing budgets and spend caps across complex hierarchies (C-suite, teams, users, agents), demanding sophisticated governance and allocation logic.

Mentioned resources

  • Stigg (Company/Product)
  • Anthropic (Company)
  • OpenAI (Company)
  • GitHub Copilot (Product)
  • OpenClaw (Product)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.