Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab
Summary
Felipe Blanes discusses the 'benchmark illusion' in building AI agents, where high performance on static benchmarks does not guarantee reliability in real-world customer use. He proposes an 'eval flywheel'—a continuous, customer-driven process—that mandates defining success based on customer needs, capturing real-world signals, diagnosing gaps (model, engineering, or product), and feeding those insights back into development decisions. The talk emphasizes that transparency about product limitations is crucial for building customer trust, noting that reliability is a binary 'trust' or 'no trust' state, rather than a linear scale.
Key takeaways
-
The Benchmark Illusion
3:42
Relying solely on static benchmarks or synthetic data (static evals) is insufficient because real customer use cases often involve unexpected actions that cause product failure. The solution is to close the customer loop.
-
The Eval Flywheel
7:21
The core process involves four steps: 1) Define success (based on customer needs, not internal assumptions); 2) Capture signals (using instrumentation and, critically, talking to customers); 3) Diagnose gaps (categorizing issues into model failures, engineering gaps, or product positioning issues); and 4) Feed decisions (prioritizing fixes for the science, engineering, and product teams).
-
The Trust Cliff
15:35
Reliability is not linear. Customers perceive 80% reliability as requiring too much manual monitoring and effort, whereas reaching a threshold (around 90-92%) shifts the perception to 'trustworthy' and allows for full workflow handover.
-
Transparency Builds Trust
It is more effective to be transparent about a product's limitations and what is out of scope than to focus only on high benchmarks. Openness about capabilities and limitations increases customer trust.
Technical details
-
Nova Act and Agent Building
100s
Nova Act is a service used to build browser agents, allowing automation of tasks performed within a web browser. The service progressed from a research preview (March) to general availability on AWS (December).
-
Gap Diagnosis Categories
590s
When diagnosing gaps from customer signals, issues fall into three categories: 1) Model inputs (where the model is failing, feeding into evals); 2) Engineering inputs (traditional bug fixes in the codebase); and 3) Product inputs (issues with product positioning or customer understanding of the product's scope).
-
Customer Journey Stages
726s
The customer journey progresses through four stages, each yielding different signals: 1) Use case discovery (understanding the core problem); 2) Early adopters (identifying what is possible, even if unexpected); 3) Scaling (identifying the 'hero use case' or common pattern); and 4) Product gaps (identifying future features needed).
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.