# How I Learned to Stop Worrying and Love the Sandbox — Matt Brockman, E2B

## Executive summary

This workshop provides a deep dive into the operational challenges of running large-scale, isolated code execution environments using sandboxes (specifically E2B). The session walks through deliberate failures—such as runaway processes, memory leaks, and disk filling—to demonstrate best practices for resource management, state persistence, and process cleanup. Key architectural topics include snapshotting the system state, managing sandbox lifecycles (idle timeouts), and implementing user-specific workspace mapping to handle thousands of concurrent workloads.

## Key takeaways

- Sandbox Fundamentals: Sandboxes provide isolated, user-specific execution environments, preventing code from interfering with other users. E2B aims to spin up sandboxes in less than 100 milliseconds (2:51).
- State Persistence via Snapshots: Sandboxes preserve a running system's state by snapshotting both memory and the filesystem, allowing work to pause and resume seamlessly, which is crucial for long-running agent tasks (2:51).
- Resource Leak Mitigation: Common issues include processes consuming excessive CPU (e.g., 99.7% usage), memory pressure, and filling the disk. These require identifying and terminating rogue processes (18:43, 20:30, 21:57).
- Sandbox Lifecycle Management: To manage costs and resources at scale, it is necessary to set a long runtime for a task but implement a short idle timeout, allowing the sandbox to shut down after the task completes (31:14).
- Scaling Workspaces: For large deployments, the assignment mechanism should move beyond simple round-robin assignment to maintain a persistent mapping of users to specific sandbox IDs (44:00).

## Technical details

- Process Management & Cleanup: The workshop demonstrated using Linux commands to identify and kill runaway processes (e.g., killing PID 1250) that consume excessive CPU or memory. Cleaning up orphaned processes is a common, critical maintenance task when reusing sandbox states (31:14).
- Template and Cache Architecture: A 'template' is defined as a snapshotted VM state. The system separates the template cache (read-only, used for building templates) from the runtime cache (writeable, used during active sessions) to manage file system integrity (58:00).
- Sandbox Identification and Addressing: The sandbox ID is the unique identifier for the isolated environment. E2B constructs URLs using the format: `port-sandboxID.app` (11:40).
- Workload Orchestration: For distributed workloads, the system must manage resource quotas (CPU/RAM) and track which user is assigned to which sandbox ID to prevent resource exhaustion and ensure user continuity (44:00).

## Practical implications

- When building systems that rely on remote, ephemeral compute (e.g., AI agent execution), robust resource monitoring and cleanup mechanisms are mandatory.
- Architecting state persistence requires careful handling of snapshots to ensure data integrity while minimizing resource overhead.
- For high-volume, multi-user applications, implementing dedicated user-to-sandbox mapping is necessary to prevent resource contention and improve user experience.
- The trade-off between speed (running active sandboxes) and cost (pausing/freezing sandboxes) must be managed through intelligent lifecycle policies.

## Topics

Sandboxing, Resource Management, Distributed Computing, AI Agent Workloads, System Orchestration, E2B, tinyurl.com

Source: https://www.youtube.com/watch?v=fz6-NS7qpZc
