Topic

Open-Weight Models

All digests tagged Open-Weight Models

Uber Burned 6x Its AI Budget in Four Months thumbnail

· 10:18

Uber Burned 6x Its AI Budget in Four Months

The video provides a deep dive into the operational costs and optimization challenges of building agentic AI systems. Key themes include the critical need for cache-aware routing to manage computational costs, the alarming rate of AI budget expenditure (e.g., Uber's 6x increase in four months), and the finding that only a small fraction of AI spending translates into shipped, meaningful code. Speakers advocate for leveraging open-weight models, implementing smart evaluation gates, and optimizing knowledge base updates to prevent unnecessary human intervention.

Key takeaways

  1. AI Budget Overruns are Common 3:32

    Uber increased its AI budget by six times since 2024, spending the entire increase within four months, leaving them out of budget for the remainder of the year. (03:12)

  2. Low Dollar-to-Shipped-Code Ratio 3:32

    Only $18 of every $100 spent on AI actually reaches meaningful code that gets shipped to users. (03:12)

  3. Caching is Essential for Agent Workloads 0:23

    Properly implementing caching, especially for output/input tokens, is crucial for agent workloads, as it limits computation to only newly generated tokens. (00:00:23)

  4. Open-Weight Models Handle Significant Workload 5:23

    Open-weight models running on owned hardware can now handle approximately 80% of the required work, allowing organizations to avoid vendor lock-in. (05:23)

Watch on YouTube Full article

Who’s afraid of an open-weight model? GLM, context bombing and post-Black Hat attacks thumbnail

· 26:43

Who’s afraid of an open-weight model? GLM, context bombing and post-Black Hat attacks

The discussion explores the rapid advancement and associated risks of open-weight AI models like GLM-5.3, which show strong capabilities in vulnerability discovery and validation. Defensively, researchers developed 'context bombing,' a technique using malicious prompts to shut down attacking AI agents. The conversation emphasizes that while offensive security (AI model development) is accelerating faster than defensive measures (automated patching/blue team), classic principles like defense-in-depth and assuming breach remain critical. Finally, the segment warns against sophisticated social engineering attacks targeting cybersecurity professionals post-conference.

Key takeaways

  1. AI Vulnerability Discovery is Accelerating 2:00

    Open-weight models like GLM-5.3 demonstrate advanced cyber capabilities through post-training, achieving a score of 84.5% on CyberGym for vulnerability discovery and validation, reaching parity with competitors like GPT Sol and Mythos.

  2. Context Bombing as Defensive Measure 12:10

    Tracebit researchers developed 'context bombing,' which uses malicious prompts placed alongside assets to confuse attacking AI agents. Testing showed that instances of models proceeding with an attack dropped from 91% to 15%.

  3. Blue Team Must Match Offensive Pace 4:00

    Experts stressed the need for significant investment in automated patching and blue team capabilities (e.g., automated SOC) to keep pace with AI-driven offensive security, noting that manual processes are insufficient.

Watch on YouTube Full article