AI Engineer

Agents That Write Their Own Tools at Runtime — Sandhya Subramani, AWS

Published 2026-10-04 · Duration 20:40

Summary

This talk introduces meta-tooling, a capability where AI agents can write, load, and use their own tools and even create sub-agents at runtime, without needing to be explicitly programmed for every function. Using the open-source Strands Agents SDK, the speaker demonstrates how an agent can dynamically create tools (like a math calculator or character counter) and handle complex tasks, such as planning a multi-stage itinerary. For production use, the talk emphasizes the necessity of robust guardrails, including sandboxed execution, constrained permissions, and comprehensive evaluation (evals) to ensure reliability and prevent undue damage.

Download summary

Key takeaways

  1. Meta-Tooling Core Components 10:44

    Implementing meta-tooling requires three core tools: `editor`, `shell`, and `load_tool`, guided by a system prompt that defines what constitutes a 'good tool.'

  2. Self-Modifying Agents 19:13

    Agents can write tools and even create sub-agents (e.g., a travel planner creating 'flight agent,' 'activities agent,' and 'itinerary agent') to solve complex problems.

  3. Agentic Patterns 18:23

    Advanced multi-agent systems can utilize patterns like Swarm (parallel sub-agents), Graph, Handoff, and Workflow to coordinate tasks.

  4. Production Guardrails

    To trust self-modifying agents, implement guardrails including sandboxed execution, constrained permissions, and thorough evaluation (evals) across goal success, tool choice, and inter-agent flow.

Technical details

  • Meta-Tooling Implementation 644s

    The Strands Agents SDK enables meta-tooling using a system prompt and three foundational tools: `editor`, `shell`, and `load_tool`. This allows the agent to dynamically generate and utilize custom tools at runtime.

  • Agentic Patterns 1103s

    Three primary patterns for multi-agent systems are Swarm (parallel task execution), Graph (sequential dependency), and Handoff/Workflow (combinations of the above).

  • Agent Evaluation (Evals)

    Reliability requires evaluating: 1) End goal achievement (Did it book the flight?), 2) Trace level (Was the answer helpful/correct?), 3) Tool access (Was the right tool/parameter used?), and 4) Inter-agent dependency (Was the correct sequence followed?).

  • Security and Guardrails

    Essential guardrails include sandboxing the code execution environment, constraining permissions, and maintaining observability to prevent the agent from causing undue damage.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.