Topic

AI Evaluation

All digests tagged AI Evaluation

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo thumbnail

· 19:48

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

This talk addresses the critical challenge of evaluating high-stakes AI systems, particularly ambient scribes in healthcare, where dangerous failures often manifest as subtle omissions or hallucinations rather than obvious errors. The speaker argues that traditional verification methods (like fixed rubrics or simple difference checks) fail because the 'standard of good' is tacit, contextual, and constantly evolving. A proposed solution involves building a continuous evaluation loop: Discovering failure modes from real-world outputs, capturing expert judgment on these modes, and calibrating every new output against this accumulated, case-specific context rather than a static rule set.

Key takeaways

  1. High-Stakes Failure Modes 2:09

    In clinical notes, the most dangerous failures are often subtle omissions (e.g., missing jaw pain symptoms) or hallucinations that look technically correct but are factually wrong. In large studies, nearly 1 in 20 notes carried an error serious enough to cause significant harm [1:29].

  2. Limitations of Current AI Evaluation 11:43

    Verification is only easy for the 'easy half' (e.g., spotting differences between transcript and note). The hard part is determining which difference—an omission, change, or addition—actually matters in context [7:03]. This judgment is tacit, contextual, and moving.

  3. The Continuous Evaluation Loop

    To overcome the limitations of static rubrics, the recommended approach is a continuous loop: 1) Discover failure modes from real outputs (building a 'failure mode ontology'), 2) Capture expert judgment on these modes, and 3) Calibrate every output against this accumulated, case-specific context, rather than relying on fixed weights or prompts [13:49].

Watch on YouTube Full article

From Ambient Documentation to Clinical Intelligence — Chaitanya Asawa, Abridge thumbnail

· 21:35

From Ambient Documentation to Clinical Intelligence — Chaitanya Asawa, Abridge

The talk details Abridge's evolution from solving clinical documentation burnout—a high-stakes administrative problem in healthcare—to building comprehensive clinical intelligence tools. The speaker emphasizes that all healthcare processes are downstream of the doctor-patient conversation. Technically, the core challenges involve maintaining extremely high quality and low latency in a high-stakes environment, requiring novel approaches like decomposing complex tasks into smaller models (instead of relying solely on frontier LLMs) and developing sophisticated evaluation systems using expert human judges and rubrics to address the small generator/verifier gap.

Key takeaways

  1. The Centrality of Conversation 5:50

    All administrative processes in healthcare (billing, clinical decision support, etc.) are built around the core conversation between a doctor and a patient. Abridge aims to automate this entire downstream machinery.

  2. The Productivity Paradox in Healthcare 10:20

    Unlike many industries where productivity increases lower costs, administrative costs in healthcare have continued to rise over decades, creating a significant operational burden that technology must address.

  3. High Stakes AI Development

    In clinical decision support, the cost of being wrong is extremely high. This necessitates rigorous quality control and evaluation methods far beyond typical generative AI applications.

Watch on YouTube Full article