# Uber Burned 6x Its AI Budget in Four Months

## Executive summary

The video provides a deep dive into the operational costs and optimization challenges of building agentic AI systems. Key themes include the critical need for cache-aware routing to manage computational costs, the alarming rate of AI budget expenditure (e.g., Uber's 6x increase in four months), and the finding that only a small fraction of AI spending translates into shipped, meaningful code. Speakers advocate for leveraging open-weight models, implementing smart evaluation gates, and optimizing knowledge base updates to prevent unnecessary human intervention.

## Key takeaways

- AI Budget Overruns are Common: Uber increased its AI budget by six times since 2024, spending the entire increase within four months, leaving them out of budget for the remainder of the year. (03:12)
- Low Dollar-to-Shipped-Code Ratio: Only $18 of every $100 spent on AI actually reaches meaningful code that gets shipped to users. (03:12)
- Caching is Essential for Agent Workloads: Properly implementing caching, especially for output/input tokens, is crucial for agent workloads, as it limits computation to only newly generated tokens. (00:00:23)
- Open-Weight Models Handle Significant Workload: Open-weight models running on owned hardware can now handle approximately 80% of the required work, allowing organizations to avoid vendor lock-in. (05:23)

## Technical details

- Cache-Aware Routing: For multi-replica model serving, it is critical to ensure that subsequent turns in an agent workflow land on the same replica where the cache is stored, preventing wasted computation. (00:02:06)
- Agentic Workflow Optimization: The optimization goal is to ensure that the vast majority of output/input tokens are cached, minimizing GPU work only to the new tokens being generated. (00:00:23)
- Automated Knowledge Base Updates: A robust system should use a pipeline (e.g., GitHub Action) to review knowledge base article changes, determine if the change is minor/moderate/major, and run evaluations. Human review should only be required for major changes. (07:54)

## Practical implications

- Implement cache-aware routing mechanisms when deploying agentic models across multiple replicas to minimize compute costs.
- Prioritize investing in local, open-weight models and owned hardware to reduce dependency on proprietary LLM ecosystems.
- Design CI/CD pipelines for knowledge bases that automatically run evaluations and only trigger human review for significant, complex content changes.
- Focus engineering efforts on defining clear success criteria for evaluations rather than manually reviewing text diffs on skills.

## Topics

Agentic Workflows, LLM Cost Management, Caching Strategies, Open-Weight Models, Knowledge Graph Maintenance, Build Automation, AI DevCon NYC 2026, Tessl

Source: https://www.youtube.com/watch?v=nnja7kk4EPM
