Presentation Information
[2G5-OS-47b-03]Action-Chunk Correction for VLA Policies Enabling Real-Time Responsiveness Under Inference Latency
〇Kohei Sendai1, Maxime Alvarez1, Tatsuya Matsushima1, Yutaka Matsuo1, Yusuke Iwasawa1 (1. The University of Tokyo)
Keywords:
Physical AI
Vision-Language-Action (VLA) models typically utilize action chunking to enhance efficiency and temporal coherence, yet this approach often compromises reactivity under inference delays and long execution horizons. To address this, we introduce Asynchronous Action Chunk Correction (A2C2), a lightweight, real-time correction head designed to refine action chunks from off-the-shelf VLAs without requiring retraining. A2C2 operates at every control step by integrating the latest observations, base actions, and positional features to output precise, time-aware corrections. This restores closed-loop responsiveness while maintaining the base model's competence. Evaluated on the dynamic Kinetix suite and LIBERO Spatial benchmark, A2C2 demonstrates consistent success rate improvements over baselines like Real Time Chunking (RTC), achieving gains of +23\% and +7\% points respectively under varying delays. With minimal computational overhead, A2C2 serves as an effective plug-in mechanism for deploying high-capacity VLA policies in real-time control.
