# Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab

## Executive summary

Felipe Blanes discusses the 'benchmark illusion' in building AI agents, where high performance on static benchmarks does not guarantee reliability in real-world customer use. He proposes an 'eval flywheel'—a continuous, customer-driven process—that mandates defining success based on customer needs, capturing real-world signals, diagnosing gaps (model, engineering, or product), and feeding those insights back into development decisions. The talk emphasizes that transparency about product limitations is crucial for building customer trust, noting that reliability is a binary 'trust' or 'no trust' state, rather than a linear scale.

## Key takeaways

- The Benchmark Illusion: Relying solely on static benchmarks or synthetic data (static evals) is insufficient because real customer use cases often involve unexpected actions that cause product failure. The solution is to close the customer loop.
- The Eval Flywheel: The core process involves four steps: 1) Define success (based on customer needs, not internal assumptions); 2) Capture signals (using instrumentation and, critically, talking to customers); 3) Diagnose gaps (categorizing issues into model failures, engineering gaps, or product positioning issues); and 4) Feed decisions (prioritizing fixes for the science, engineering, and product teams).
- The Trust Cliff: Reliability is not linear. Customers perceive 80% reliability as requiring too much manual monitoring and effort, whereas reaching a threshold (around 90-92%) shifts the perception to 'trustworthy' and allows for full workflow handover.
- Transparency Builds Trust: It is more effective to be transparent about a product's limitations and what is out of scope than to focus only on high benchmarks. Openness about capabilities and limitations increases customer trust.

## Technical details

- Nova Act and Agent Building: Nova Act is a service used to build browser agents, allowing automation of tasks performed within a web browser. The service progressed from a research preview (March) to general availability on AWS (December).
- Gap Diagnosis Categories: When diagnosing gaps from customer signals, issues fall into three categories: 1) Model inputs (where the model is failing, feeding into evals); 2) Engineering inputs (traditional bug fixes in the codebase); and 3) Product inputs (issues with product positioning or customer understanding of the product's scope).
- Customer Journey Stages: The customer journey progresses through four stages, each yielding different signals: 1) Use case discovery (understanding the core problem); 2) Early adopters (identifying what is possible, even if unexpected); 3) Scaling (identifying the 'hero use case' or common pattern); and 4) Product gaps (identifying future features needed).

## Practical implications

- Shift testing focus from internal, synthetic benchmarks to real-world customer usage patterns.
- Implement a continuous feedback loop (eval flywheel) that prioritizes customer-reported needs over internal assumptions.
- When reporting reliability, frame the discussion around the 'trust cliff' concept, recognizing that reliability is a binary state for the end-user.
- Ensure product documentation and customer communication are transparent about current limitations and out-of-scope functionality.

## Topics

AI Agents, Reliability Engineering, Customer Experience (CX), Product Development Lifecycle, AWS Services, Amazon Nova Act, Nova Act docs

Source: https://www.youtube.com/watch?v=Emo5FGGY-wM
