# Why AI Agents Should Have Their Own Sandbox — Philipp Schmid, Google DeepMind

## Executive summary

Philipp Schmid details the shift from traditional LLM interactions (user/model turns) to sophisticated, managed AI agents. The core architectural change is the Gemini Interactions API, which provides server-side state management, long-running background jobs, and a structured 'steps' timeline. Crucially, agents now operate within a managed cloud sandbox, allowing them to run code, install dependencies, and share persistent environments without the user managing underlying infrastructure. This capability is packaged via the Agents API, enabling teams to create and share complex, multi-step agents easily.

## Key takeaways

- Interactions API for Agents and Models: The Gemini Interactions API unifies the interface for both models and agents, replacing the older Responses API. It supports multimodal inputs (text, audio, images) and provides server-side state management and long-running operations (e.g., video generation), simplifying complex application development.
- Agentic Workflow: From Turns to Steps: Agentic applications move away from the simple user-model turn-based conversation model. Instead, they utilize a synchronous timeline of 'steps,' where every action or generation is explicitly defined, making the system easier to extend and reason about.
- Managed Sandboxes and Environments: Agents run in a managed cloud sandbox (the 'environment' field), which handles code execution, file system operations, and dependency installation. This environment is persistent and can be shared between multiple agents or sub-agents, eliminating the need for developers to manage infrastructure like Kubernetes or Terraform.
- Agent Packaging and Sharing: The Agents API allows developers to define and package base agents, including instructions and environments. This enables teams to create reusable agents and share them with users without writing complex infrastructure code.

## Technical details

- Gemini Interactions API: The API supports server-side state management and long-running operations (async requests). It handles multimodal inputs and simplifies tool use by allowing the model to decide whether to use built-in tools (like Google Search) or custom functions.
- Agent Architecture: The system moves from a user/model turn concept to a structured timeline of 'steps.' This change is critical because function calling results are now treated as environment outputs, not user inputs.
- Agent Components: An agent is defined by its underlying 'harness' (the backend logic), its system instructions (often loaded from an `AGENTS.md` file), and its environment. The environment provides the sandbox for running code and managing state.
- Development Tools: Developers can use the Gemini API CLI to interact with the API and manage agents. The platform also supports defining agents using the Agents API, allowing the creation of environments that can then be used to instantiate new, shareable agents.

## Practical implications

- Build engineers can now design complex, multi-step agent workflows without managing the underlying cloud infrastructure or state persistence.
- The standardized Agents API simplifies the creation of reusable, shareable AI components, accelerating product development cycles.
- The shift to a 'steps' timeline provides a more robust and predictable model for debugging and extending agent capabilities compared to traditional turn-based APIs.

## Topics

Generative AI, Large Language Models (LLMs), Agentic Systems, API Design, Cloud Computing, State Management, Gemini Interactions API, Gemini API agents, Gemini API CLI

Source: https://www.youtube.com/watch?v=oWTEiYpxl80
