# The $600 PR: Where Software Factory ROI Really Comes From

## Executive summary

The true Return on Investment (ROI) in an AI software factory is not derived from the coding agent or model itself, but from the surrounding system components: context management, verification layers, and robust feedback loops. The presentation outlines a five-stage factory model, emphasizing that optimizing the system's ability to learn from failures and improve its inputs (context, environment) yields greater returns than optimizing the agent's output.

## Key takeaways

- ROI is system-centric, not agent-centric: The majority of ROI comes from improving the processes surrounding the agent (context, verification, feedback loops), rather than solely improving the agent or model itself. (00:02:56)
- The five stages of a software factory: The process involves defining a useful task, agent execution in a dev sandbox, verification/delivery, and critically, establishing a feedback loop to improve the system overall. (00:02:56)
- Focus on behavioral, not code-level, human oversight: Human intervention should focus on defining desired system behavior (e.g., 'we want this kind of behavior expressed') rather than reviewing every line of code. (00:04:42)
- Full cost accounting is crucial for accurate ROI: When measuring ROI, the full cost must include retries, reviews, infrastructure, and human attention, which are often far more expensive than token costs. (00:21:38)
- Context and environment are key inputs: Feeding the agent context from external sources (Slack, Notion docs, incident reports) and ensuring the environment can stand up mocks/deps is critical for accurate verification. (00:08:55)

## Technical details

- Agentic Workflow Failure Example: A PR update cycle involving two agents resulted in a $600 burn because the agents executed exactly what they were told, lacking built-in assumptions like 'this is far enough' or 'the scope is exactly targeting this.' (00:01:31)
- Formal Verification: Using tools like Quint, formal modeling was applied to a small, buggy subsystem (e.g., a durable work queue) to verify its behavior and knock out an entire class of bugs, providing stability for downstream components. (00:12:37)
- Metrics for ROI: Beyond simple CI pass rates, measuring ROI requires analyzing the full cost, including first review approval rate and the cost of review rounds, to understand where optimization efforts yield the greatest savings. (00:14:52)
- Observability at the Meta Layer: Instead of focusing on agent transcripts, better signals are derived by correlating data across time series (e.g., linking production incidents back to specific PR review cycles or missing verification functions). (00:21:38)
- Merge Queue Analysis: Analyzing the merge queue (P90 at 52.4 minutes) revealed that the primary bottleneck was waiting time, not CI duration (P90 at 15.5 minutes), highlighting the importance of specifying the correct system constraint. (00:21:38)

## Practical implications

- Shift focus from optimizing agent prompts/models to optimizing the surrounding system infrastructure (context ingestion, verification, feedback loops).
- Implement full cost accounting for engineering work, factoring in human attention and review cycles alongside token costs.
- Use formal verification (e.g., Quint) on critical, small subsystems to guarantee invariants and eliminate entire classes of bugs.
- Improve agent interaction by asking poorly specified questions in public forums (like Slack) to generate and refine context, rather than assuming the agent can read your mind.

## Topics

AI Software Factory, Build Engineering, DevOps, Formal Verification, LLM Orchestration, ROI Measurement, Tessl, Quint, Anthropic, OpenAI, Google

Source: https://www.youtube.com/watch?v=59TBijDBbQg
