AI Engineer

Why Bigger Context Windows Won't Save Your Agent — Elizabeth Fuentes Leone, AWS

Published 2026-10-08 · Duration 15:54

Summary

This talk addresses the limitations of large language model (LLM) agents when processing massive amounts of data, such as extensive application logs or tool outputs. The core finding is that simply increasing the context window size is insufficient because models suffer from an 'attention curve' problem, forgetting information in the middle of a long context. The solution, termed 'Context Engineering,' involves four strategies—Externalize, Select, Compress, and Isolate—to manage context optimally. Key techniques include using memory pointers for large data storage, implementing structured memory (short-term, long-term, graph), and managing agent state to prevent endless tool loops and handle slow external API calls (MCP tools) using asynchronous handles.

Download summary

Key takeaways

  1. Bigger Context Windows Are Not the Fix 2:37

    Models tend to remember the beginning and end of a long context but lose information in the middle, a phenomenon described by the attention curve. Context engineering is required to provide the model only with the necessary information for optimal outcomes and token consumption. (1:57, 2:12)

  2. The Four Context Engineering Strategies 5:17

    To manage context, implement strategies: 1) Externalize (move big data to persistent storage via memory pointers), 2) Select (retrieve only relevant information), 3) Compress (summarize or compact data), and 4) Isolate (separate context across different agents). (3:17)

  3. Advanced Memory Management 10:42

    Agents should utilize a combination of memory types: Short-term (recent conversation history), Long-term (stored in a vector database), and Graph memory (using entity graphs to understand relationships). (6:42)

  4. Handling Large Tool Outputs 13:32

    Instead of passing large logs directly, use a 'memory pointer' tool. The tool saves the data to external storage and only passes the unique ID (pointer) to the agent's context window. (8:12)

  5. Preventing Agent Loops and Delays

    To prevent agents from entering infinite loops, set limits on tool calls (e.g., `max tool counts`). For slow external APIs (MCP tools), use asynchronous handles to prevent the application from timing out while waiting for a response. (11:22, 12:12)

Technical details

  • Strands Agents Framework 507s

    Strands Agents is an open-source, model-agnostic framework that simplifies agent loop creation, allowing agents to manage context using built-in strategies like sliding window and summarization conversation managers. (5:07)

  • Conversation Managers 597s

    Conversation managers can summarize older messages while preserving the most recent messages (e.g., summarizing 50% of history while keeping the last four messages). (5:57)

  • Multi-Agent State Sharing 947s

    When using a 'swarm' of agents, sharing context via a memory pointer ID (rather than passing all information) prevents context pollution and keeps the state manageable. (9:47)

  • Anti-Patterns to Avoid

    Avoid 'context stuffing' (passing all available data), relying solely on native summarization (which may omit critical details), and general context population (passing unnecessary information). (13:41)

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.