# AI ROI, Why the Agent Isn't the Answer

## Executive summary

Achieving Return on Investment (ROI) with AI agents requires moving beyond simply implementing agents and instead focusing on the surrounding system architecture. Key constraints include the quality of the data context (what the agent can see) and the ability to measure performance (how to prove it's working). Speakers emphasized that the most valuable investments are in creating robust feedback loops, formal verification, and building federated, comprehensive data layers that preserve optionality and scope.

## Key takeaways

- Focus on the Feedback Loop, Not Just the Agent: The greatest ROI comes from investing in the feedback loop—the ability to automatically ingest failure modes and allow the system to improve itself. This is more critical than optimizing the agent itself.
- Measure Full System Cost, Not Just Token Cost: To accurately measure ROI, the cost model must include all factors: retries, human attention, review cycles, and infrastructure, not just the token expenditure. Analyzing the full system cost allows for better optimization decisions.
- Data Context is the Primary Constraint: Agents are limited by the data they receive. The solution is to build a federated, logical view of data that stitches together multiple sources (e.g., real-time Kafka data and long-term Iceberg data) to preserve optionality and scope.
- Formal Verification for Stability: For mission-critical components, formal modeling (e.g., using TLA) is necessary to verify system behavior and prevent bugs, even when agents are constantly modifying the code base.

## Technical details

- Agentic Workflow Architecture: A robust agentic workflow involves five stages: defining a useful task, establishing shared context, agent execution (in a dev sandbox), verification/delivery, and continuous feedback loop improvement.
- Data Blind Spots in AI: Common data blind spots include outdated data, wrong scoping, missing relationships between departments/products, and the inability to distinguish between 'no data' and 'unknown state'.
- Data Federation and Logical Views: Instead of relying on single data sources, the best practice is to create a logical view by federating across multiple systems (e.g., combining real-time data from Apache Kafka with historical data in Apache Iceberg) to provide a comprehensive context.
- Formal Modeling: Formal methods (like TLA) can be used to model and verify complex system behaviors (e.g., a work queue/Q) to ensure stability and prevent bugs introduced by autonomous agents.

## Practical implications

- Shift focus from optimizing the agent's output to optimizing the data context and the system's feedback mechanisms.
- Implement comprehensive observability at the meta-layer (e.g., tracking queue wait times vs. CI run times) to identify bottlenecks outside the agent.
- Adopt a data federation approach to build logical views, combining real-time and historical data sources to give agents a complete, consistent context.
- When measuring ROI, ensure the cost model accounts for human time, review cycles, and retries, not just token usage.

## Topics

Build Engineering, AI Architecture, Data Engineering, CI/CD, System Reliability, Anthropic, OpenAI, Google Cloud, TLA (Temporal Logic of Actions), Apache Kafka, Apache Iceberg, Open Metadata

Source: https://www.youtube.com/watch?v=LZ-1khXvnrI
