Why Bigger Context Windows Won't Save Your Agent — Elizabeth Fuentes Leone, AWS
This talk addresses the limitations of large language model (LLM) agents when processing massive amounts of data, such as extensive application logs or tool outputs. The core finding is that simply increasing the context window size is insufficient because models suffer from an 'attention curve' problem, forgetting information in the middle of a long context. The solution, termed 'Context Engineering,' involves four strategies—Externalize, Select, Compress, and Isolate—to manage context optimally. Key techniques include using memory pointers for large data storage, implementing structured memory (short-term, long-term, graph), and managing agent state to prevent endless tool loops and handle slow external API calls (MCP tools) using asynchronous handles.
Key takeaways
-
Bigger Context Windows Are Not the Fix
2:37
Models tend to remember the beginning and end of a long context but lose information in the middle, a phenomenon described by the attention curve. Context engineering is required to provide the model only with the necessary information for optimal outcomes and token consumption. (1:57, 2:12)
-
The Four Context Engineering Strategies
5:17
To manage context, implement strategies: 1) Externalize (move big data to persistent storage via memory pointers), 2) Select (retrieve only relevant information), 3) Compress (summarize or compact data), and 4) Isolate (separate context across different agents). (3:17)
-
Advanced Memory Management
10:42
Agents should utilize a combination of memory types: Short-term (recent conversation history), Long-term (stored in a vector database), and Graph memory (using entity graphs to understand relationships). (6:42)
-
Handling Large Tool Outputs
13:32
Instead of passing large logs directly, use a 'memory pointer' tool. The tool saves the data to external storage and only passes the unique ID (pointer) to the agent's context window. (8:12)
-
Preventing Agent Loops and Delays
To prevent agents from entering infinite loops, set limits on tool calls (e.g., `max tool counts`). For slow external APIs (MCP tools), use asynchronous handles to prevent the application from timing out while waiting for a response. (11:22, 12:12)