Presentation Information
[2G5-OS-47b-01]Multi-modal Navigation Model
Shintaro Nakaoka1, 〇Shigemichi Matsuzaki1, Kazuhito Tanaka1, Noriaki Hirose2,3 (1. Toyota Motor Corporation, 2. Toyota Motor North America, 3. University of California, Berkeley)
Keywords:
robot navigation,foundation models,multi-modal models
This paper presents a multi-modal foundation model for autonomous navigation that jointly processes RGB images, depth images, and 2D LiDAR scans for better performance, generalization, and flexibility. By extending a Transformer-based visual navigation backbone with multi-modal tokenization and random modality masking, the model can be trained with heterogeneous datasets and deployed on robots with arbitrary sensor configurations. We present the overview of the ongoing research, future experiment and development plans.
