Presentation Information

[2K6-GS-7d-03]Visual Access Boundaries in Vision-Language Models

〇Hiroto Osaka1, Shohei Taniguchi1, Gouki Minegishi1, Kai Yamashita1, Masahiro Suzuki1, Yutaka Matsuo1 (1. University of Tokyo)

Keywords:

Vision-Language Model,Chain-of-Thought,Perception–Reasoning Decomposition,Dynamic Visual Gating,Causal Intervention

Chain-of-Thought (CoT) prompting has emerged as a powerful test-time scaling strategy for Vision-Language Models (VLMs).
However, it remains unclear whether longer reasoning extends visual feature extraction or merely amplifies language-side computation.
We introduce the Visual Access Boundary (VAB), defined via a causal intervention that systematically blocks attention to image tokens across layer depth and generation time.
Across models and tasks, we find that CoT consistently shifts the VAB to deeper layers (deep shift), despite dramatically increasing generation length.
This indicates that CoT does not proportionally prolong visual access over time, but instead increases the depth required for stable perceptual grounding.
The magnitude of this shift is modest relative to token growth, suggesting that CoT primarily expands language-space reasoning conditioned on perceptually grounded symbols.
These findings provide a mechanistic account of how perception and reasoning interact in VLMs.