AI Engineer

Brains vs Hands: How to Run AI Agents Safely in Production — Viren Baraiya

Published 2026-10-08 · Duration 16:50

Summary

Running AI agents in production requires a fundamental architectural shift: the agent's role must be separated from the execution mechanism. The speaker emphasizes that the LLM (the 'brain') should only be responsible for planning the next steps, while a deterministic, durable workflow engine (the 'hands' or harness) must handle the actual execution. This approach treats agent harnesses as 'late-bound sagas,' ensuring reliability, managing side effects, and guaranteeing deterministic outcomes, which is crucial for mission-critical systems like SRE or payment processing.

Download summary

Key takeaways

  1. Agent Scope Beyond Chatbots 0:01

    In production, agents are not limited to simple chatbots. They must handle background tasks, run on schedules, react to events (e.g., logs, alerts), and coordinate in multi-agent systems (1:52, 2:27).

  2. Harness as the Application 0:03

    An agent, when combined with a controlling harness, functions as an application, much like a set of microservices. The harness is responsible for delivering the overall business goal, integrating databases, internal systems, and human-in-the-loop approvals (3:02, 3:12).

  3. The Brain and the Hands Separation 0:09

    The core principle is the clear split: the LLM plans what should happen next (non-deterministic), but the harness executes the plan using deterministic code. This ensures reliable execution, especially for critical tasks like cluster restarts (9:56).

  4. Agentic Workflows as Late-Bound Sagas 0:10

    Agent harnesses are described as 'late-bound sagas.' They gain the benefits of traditional sagas (visibility, control) while allowing the agent to propose and build the workflow at runtime, rather than requiring the entire sequence to be defined upfront (10:21).

Technical details

  • Agentic Architecture 5s

    The system must combine both deterministic (e.g., workflow execution) and non-deterministic (LLM reasoning) parts. The harness must guarantee deterministic behavior for critical steps, such as cluster restarts, and must support idempotency and side-effect recording (5:02, 9:41).

  • Durability and State Management 7s

    Because harnesses are long-running processes (potentially running for days or months), durability is a mandatory requirement. The system must maintain a state of the world, recording all completed work and side effects to allow for recovery from failures (6:41, 7:26).

  • Workflow Execution 11s

    The process involves the LLM generating a plan, which is then compiled into a fully runnable, deterministic workflow using an orchestration engine like Conductor. This allows the system to execute complex, multi-step, self-planning loops (11:51).

Mentioned resources

  • Orkes (Company/Product)
  • Conductor (Workflow Orchestration Platform)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.