AI Native Dev

Uber Burned 6x Its AI Budget in Four Months

Published 2026-09-28 · Duration 10:18

Summary

The video provides a deep dive into the operational costs and optimization challenges of building agentic AI systems. Key themes include the critical need for cache-aware routing to manage computational costs, the alarming rate of AI budget expenditure (e.g., Uber's 6x increase in four months), and the finding that only a small fraction of AI spending translates into shipped, meaningful code. Speakers advocate for leveraging open-weight models, implementing smart evaluation gates, and optimizing knowledge base updates to prevent unnecessary human intervention.

Download summary

Key takeaways

  1. AI Budget Overruns are Common 3:32

    Uber increased its AI budget by six times since 2024, spending the entire increase within four months, leaving them out of budget for the remainder of the year. (03:12)

  2. Low Dollar-to-Shipped-Code Ratio 3:32

    Only $18 of every $100 spent on AI actually reaches meaningful code that gets shipped to users. (03:12)

  3. Caching is Essential for Agent Workloads 0:23

    Properly implementing caching, especially for output/input tokens, is crucial for agent workloads, as it limits computation to only newly generated tokens. (00:00:23)

  4. Open-Weight Models Handle Significant Workload 5:23

    Open-weight models running on owned hardware can now handle approximately 80% of the required work, allowing organizations to avoid vendor lock-in. (05:23)

Technical details

  • Cache-Aware Routing 126s

    For multi-replica model serving, it is critical to ensure that subsequent turns in an agent workflow land on the same replica where the cache is stored, preventing wasted computation. (00:02:06)

  • Agentic Workflow Optimization 23s

    The optimization goal is to ensure that the vast majority of output/input tokens are cached, minimizing GPU work only to the new tokens being generated. (00:00:23)

  • Automated Knowledge Base Updates 474s

    A robust system should use a pipeline (e.g., GitHub Action) to review knowledge base article changes, determine if the change is minor/moderate/major, and run evaluations. Human review should only be required for major changes. (07:54)

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.