Topic

User Experience (UX)

All digests tagged User Experience (UX)

Act, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku thumbnail

· 20:25

Act, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku

The presentation argues that improving voice AI user experience requires focusing on a second, often neglected dimension: system behavior under uncertainty. While increasing accuracy (Knob One) is critical, the system's ability to intelligently decide what to do when it is unsure (Knob Two) can yield greater user satisfaction. This is quantified using the Outcome User Cost Heuristic (OUCH), which minimizes the total user effort by assigning differential costs to various bad outcomes (e.g., playing the wrong song vs. simply stating 'I did not understand').

Key takeaways

  1. The Two Knobs of Voice AI Improvement 0:03

    User satisfaction can be improved by increasing technical accuracy (Knob One) or by optimizing the system's decision-making process when confidence is low (Knob Two). The latter is often overlooked.

  2. The Outcome User Cost Heuristic (OUCH) 0:10

    Instead of treating all errors equally, OUCH minimizes the total user cost by quantifying the relative pain of different bad outcomes (e.g., the effort required to stop a wrong song vs. the time taken to hear 'Sorry, I did not understand').

  3. Adding Conversational Behavior 0:13

    Introducing a third behavior—confirming the guess out loud (e.g., 'Did you mean ABC?')—splits the confidence range into three regions (Stop, Confirm, Act) and further lowers the overall user cost.

Watch on YouTube Full article

"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow thumbnail

· 15:05

"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow

Midam Kim presents a linguistic framework for diagnosing failures in voice AI, arguing that these failures are not isolated bugs but structured issues. She proposes that human communication is a 'joint activity' involving the continuous updating of a 'mental model.' The framework maps this process onto two channels (listening and speaking) and four interdependent levels: sounds, words, interaction, and mental model. Successful voice AI requires holistic orchestration across all these layers, rather than optimizing components (like ASR or TTS) in isolation.

Key takeaways

  1. Voice AI is a Joint Activity 5:00

    Human communication is a joint activity where both parties contribute sounds and words, continuously updating a shared mental model. Voice AI systems must replicate this joint nature to be effective.

  2. The Linguistic Framework 11:54

    The system must be analyzed across two channels (listening/speaking) and four interdependent levels: sounds, words, interaction, and mental model. Failure in one area (e.g., STT failure at the sound level) impacts the entire system.

  3. Mental Model Accumulation

    Unlike text chat where history remains visible, in voice interactions, sounds and words vanish. The only persistent element that matters for user satisfaction is the user's accumulating mental model.

  4. System Adaptability is Key

    The system must be designed to be dynamic, adapting to context, emotion, and language change over the course of the call, rather than functioning as a static pipeline.

Watch on YouTube Full article