Presentation Information
[4F5-OS-29c-06]Vision-Language Model-Guided Dual Process Reasoning for Cooperative Multi-Agent Intention Inference
〇Kazuki Osamura1, Hidetsugu Uchida1, Chikako Matsumoto1, Narishige Abe1 (1. Fujitsu Ltd.)
Keywords:
Cooperative Multi-Agent Systems,Dual-Process Reasoning,Vision-Language Models
In cooperative multi-agent control, it is essential to react rapidly while anticipating the intentions of other agents. Conventional Overcooked agents rely on LLM-based reasoning over symbolic or text-based state representations, which limits their ability to leverage spatial cues contained in visual observations. While Vision–Language Models (VLMs) enable context understanding grounded in visual information, they are not well suited for sequential decision making due to inference latency and temporal inconsistency. In this work, we propose a dual-process cooperative agent, VLM-DPT-Agent, which extends text-based reasoning to vision-based VLM inference. The VLM estimates structured representations, including Scene Graphs, Causal Graphs, and BDI intentions, and integrates them with lightweight control mechanisms through confidence-guided intervention and a reflection module, enabling fast and stable cooperative behaviors. Experiments on the Overcooked benchmark demonstrate that our approach achieves a maximum score of 150 and consistently improves cooperative performance.
