Why AI Agents Should Have Their Own Sandbox — Philipp Schmid, Google DeepMind
Philipp Schmid details the shift from traditional LLM interactions (user/model turns) to sophisticated, managed AI agents. The core architectural change is the Gemini Interactions API, which provides server-side state management, long-running background jobs, and a structured 'steps' timeline. Crucially, agents now operate within a managed cloud sandbox, allowing them to run code, install dependencies, and share persistent environments without the user managing underlying infrastructure. This capability is packaged via the Agents API, enabling teams to create and share complex, multi-step agents easily.
Key takeaways
-
Interactions API for Agents and Models
5:21
The Gemini Interactions API unifies the interface for both models and agents, replacing the older Responses API. It supports multimodal inputs (text, audio, images) and provides server-side state management and long-running operations (e.g., video generation), simplifying complex application development.
-
Agentic Workflow: From Turns to Steps
10:45
Agentic applications move away from the simple user-model turn-based conversation model. Instead, they utilize a synchronous timeline of 'steps,' where every action or generation is explicitly defined, making the system easier to extend and reason about.
-
Managed Sandboxes and Environments
15:14
Agents run in a managed cloud sandbox (the 'environment' field), which handles code execution, file system operations, and dependency installation. This environment is persistent and can be shared between multiple agents or sub-agents, eliminating the need for developers to manage infrastructure like Kubernetes or Terraform.
-
Agent Packaging and Sharing
18:40
The Agents API allows developers to define and package base agents, including instructions and environments. This enables teams to create reusable agents and share them with users without writing complex infrastructure code.