Why AI Agents Should Have Their Own Sandbox — Philipp Schmid, Google DeepMind
Summary
Philipp Schmid details the shift from traditional LLM interactions (user/model turns) to sophisticated, managed AI agents. The core architectural change is the Gemini Interactions API, which provides server-side state management, long-running background jobs, and a structured 'steps' timeline. Crucially, agents now operate within a managed cloud sandbox, allowing them to run code, install dependencies, and share persistent environments without the user managing underlying infrastructure. This capability is packaged via the Agents API, enabling teams to create and share complex, multi-step agents easily.
Key takeaways
-
Interactions API for Agents and Models
5:21
The Gemini Interactions API unifies the interface for both models and agents, replacing the older Responses API. It supports multimodal inputs (text, audio, images) and provides server-side state management and long-running operations (e.g., video generation), simplifying complex application development.
-
Agentic Workflow: From Turns to Steps
10:45
Agentic applications move away from the simple user-model turn-based conversation model. Instead, they utilize a synchronous timeline of 'steps,' where every action or generation is explicitly defined, making the system easier to extend and reason about.
-
Managed Sandboxes and Environments
15:14
Agents run in a managed cloud sandbox (the 'environment' field), which handles code execution, file system operations, and dependency installation. This environment is persistent and can be shared between multiple agents or sub-agents, eliminating the need for developers to manage infrastructure like Kubernetes or Terraform.
-
Agent Packaging and Sharing
18:40
The Agents API allows developers to define and package base agents, including instructions and environments. This enables teams to create reusable agents and share them with users without writing complex infrastructure code.
Technical details
-
Gemini Interactions API
321s
The API supports server-side state management and long-running operations (async requests). It handles multimodal inputs and simplifies tool use by allowing the model to decide whether to use built-in tools (like Google Search) or custom functions.
-
Agent Architecture
645s
The system moves from a user/model turn concept to a structured timeline of 'steps.' This change is critical because function calling results are now treated as environment outputs, not user inputs.
-
Agent Components
914s
An agent is defined by its underlying 'harness' (the backend logic), its system instructions (often loaded from an `AGENTS.md` file), and its environment. The environment provides the sandbox for running code and managing state.
-
Development Tools
1120s
Developers can use the Gemini API CLI to interact with the API and manage agents. The platform also supports defining agents using the Agents API, allowing the creation of environments that can then be used to instantiate new, shareable agents.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.