# Why Bigger Context Windows Won't Save Your Agent — Elizabeth Fuentes Leone, AWS

## Executive summary

This talk addresses the limitations of large language model (LLM) agents when processing massive amounts of data, such as extensive application logs or tool outputs. The core finding is that simply increasing the context window size is insufficient because models suffer from an 'attention curve' problem, forgetting information in the middle of a long context. The solution, termed 'Context Engineering,' involves four strategies—Externalize, Select, Compress, and Isolate—to manage context optimally. Key techniques include using memory pointers for large data storage, implementing structured memory (short-term, long-term, graph), and managing agent state to prevent endless tool loops and handle slow external API calls (MCP tools) using asynchronous handles.

## Key takeaways

- Bigger Context Windows Are Not the Fix: Models tend to remember the beginning and end of a long context but lose information in the middle, a phenomenon described by the attention curve. Context engineering is required to provide the model only with the necessary information for optimal outcomes and token consumption. (1:57, 2:12)
- The Four Context Engineering Strategies: To manage context, implement strategies: 1) Externalize (move big data to persistent storage via memory pointers), 2) Select (retrieve only relevant information), 3) Compress (summarize or compact data), and 4) Isolate (separate context across different agents). (3:17)
- Advanced Memory Management: Agents should utilize a combination of memory types: Short-term (recent conversation history), Long-term (stored in a vector database), and Graph memory (using entity graphs to understand relationships). (6:42)
- Handling Large Tool Outputs: Instead of passing large logs directly, use a 'memory pointer' tool. The tool saves the data to external storage and only passes the unique ID (pointer) to the agent's context window. (8:12)
- Preventing Agent Loops and Delays: To prevent agents from entering infinite loops, set limits on tool calls (e.g., `max tool counts`). For slow external APIs (MCP tools), use asynchronous handles to prevent the application from timing out while waiting for a response. (11:22, 12:12)

## Technical details

- Strands Agents Framework: Strands Agents is an open-source, model-agnostic framework that simplifies agent loop creation, allowing agents to manage context using built-in strategies like sliding window and summarization conversation managers. (5:07)
- Conversation Managers: Conversation managers can summarize older messages while preserving the most recent messages (e.g., summarizing 50% of history while keeping the last four messages). (5:57)
- Multi-Agent State Sharing: When using a 'swarm' of agents, sharing context via a memory pointer ID (rather than passing all information) prevents context pollution and keeps the state manageable. (9:47)
- Anti-Patterns to Avoid: Avoid 'context stuffing' (passing all available data), relying solely on native summarization (which may omit critical details), and general context population (passing unnecessary information). (13:41)

## Practical implications

- When building CI/CD agents that analyze logs or large datasets, implement external storage and memory pointers instead of relying on the agent's context window.
- Design agent workflows to use structured memory (short-term, long-term, graph) to maintain state across sessions.
- Implement guardrails on tool usage, such as setting maximum tool invocation counts, to prevent infinite loops in automated processes.
- For external API calls, utilize asynchronous handles to ensure the agent does not stall or fail due to network latency.

## Topics

LLM Agents, Context Engineering, Memory Management, Tool Calling, Build Automation, AWS, Strands Agents, Strands Agents on GitHub

Source: https://www.youtube.com/watch?v=DrfyORO8RqA
