Multimodal Collaborative Agents for Next-Gen Commerce — Nidhi Kaushik Vyas, Google DeepMind
The talk outlines a framework for multimodal collaborative agents designed to handle 'fuzzy intent' in commerce and consumer verticals. Instead of acting as simple search bar wrappers that assume well-defined user goals, these advanced agents proactively guide users who arrive with only a 'vibe.' The core mechanism is a three-stage loop—Discovery, Research, and Response—which systematically builds a working state from multimodal inputs (images, context) to determine the optimal next question or presentation format.
Key takeaways
-
Handling Fuzzy Intent
1:48
Agents must address the 'articulation gap,' recognizing that users often arrive with vague preferences rather than precise keywords. The agent's role is to proactively elicit and refine these fuzzy intents.
-
The Collaborative Loop
3:23
The system operates in a loop: Discovery (building the working state), Research (determining the best way to ask/find information), and Response (adapting the output format).
-
Prioritizing Information Gain
13:41
The agent must calculate which unknown variable, when queried, will yield the 'maximal information gain' to move the conversation forward efficiently (e.g., determining room width is critical before recommending furniture).
-
Multimodal Elicitation
5:20
For subjective preferences, visual inspiration boards and multimodal inputs are significantly more effective than text-based questioning for establishing a common language between the user and the system.