AI Engineer

"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow

Published 2026-09-15 · Duration 15:05

Summary

Midam Kim presents a linguistic framework for diagnosing failures in voice AI, arguing that these failures are not isolated bugs but structured issues. She proposes that human communication is a 'joint activity' involving the continuous updating of a 'mental model.' The framework maps this process onto two channels (listening and speaking) and four interdependent levels: sounds, words, interaction, and mental model. Successful voice AI requires holistic orchestration across all these layers, rather than optimizing components (like ASR or TTS) in isolation.

Download summary

Key takeaways

  1. Voice AI is a Joint Activity 5:00

    Human communication is a joint activity where both parties contribute sounds and words, continuously updating a shared mental model. Voice AI systems must replicate this joint nature to be effective.

  2. The Linguistic Framework 11:54

    The system must be analyzed across two channels (listening/speaking) and four interdependent levels: sounds, words, interaction, and mental model. Failure in one area (e.g., STT failure at the sound level) impacts the entire system.

  3. Mental Model Accumulation

    Unlike text chat where history remains visible, in voice interactions, sounds and words vanish. The only persistent element that matters for user satisfaction is the user's accumulating mental model.

  4. System Adaptability is Key

    The system must be designed to be dynamic, adapting to context, emotion, and language change over the course of the call, rather than functioning as a static pipeline.

Technical details

  • Linguistic Framework Components 714s

    The framework consists of two channels (listening and speaking) and four levels: 1) Sounds (e.g., ASR/STT accuracy); 2) Words (understanding vocabulary); 3) Interaction (turn-taking/timing); and 4) Mental Model (tracking user intent). These components are interdependent; improving sounds requires considering the word level, and so on.

  • Voice AI Failure Modes 0s

    Observed failures include: 1) Misrecognition (e.g., bot hearing 'Midan' instead of 'Midam'); 2) STT failure (cutting off the user while reading an account number); 3) Lack of context retention (asking for repeated information without tracking the mental model); and 4) Failure to initiate interactive clarification.

  • System Optimization Focus

    Improvements must go beyond optimizing individual components (ASR/TTS). Focus areas include: good turn-to-turn detection, latency management, emotion detection, context retention, and dynamic orchestration.

Mentioned resources

  • Eva bench (Benchmark)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.