2606.06076v1 Jun 04, 2026 cs.AI

기호 상태로부터 시각 공간 계획 학습: 모달리티 격차 인식 자기 증류

Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation

Zhizhou Zhong
Zhizhou Zhong
Citations: 258
h-index: 4
Quan Shi
Quan Shi
Citations: 307
h-index: 4
Haochen Luo
Haochen Luo
Citations: 134
h-index: 4
Xiu Li
Xiu Li
Citations: 155
h-index: 6
Jiahui Liu
Jiahui Liu
Citations: 57
h-index: 2
Ruicheng Zhang
Ruicheng Zhang
Citations: 53
h-index: 4
Jiaqi Huang
Jiaqi Huang
Citations: 7
h-index: 2
Zunnan Xu
Zunnan Xu
Tsinghua University
Citations: 2,003
h-index: 15
Jun Zhou
Jun Zhou
Citations: 144
h-index: 5

비전-언어 모델은 일반적인 다중 모드 이해에 뛰어난 성능을 보이지만, 여전히 시각 공간 계획에는 어려움을 겪습니다. 이는 인지-추론 모달리티 격차 때문입니다. 시각적 계획은 모델이 픽셀로부터 잠재적인 상태 구조를 추론하고, 회복된 구조를 바탕으로 유효한 행동을 생성해야 하지만, 기호적 계획은 명시적인 객체와 제약을 직접 활용합니다. 이러한 차이는 시각적 상태 복구 및 다단계 계획 과정에서 이중적인 병목 현상을 야기합니다. 이를 해결하기 위해, 우리는 모달리티 격차를 인식하는 두 단계의 자기 증류 프레임워크인 MGSD를 제안합니다. 첫 번째 단계에서는 초기 학습 단계를 통해 시각적 모델에 신뢰할 수 있는 상태 표현을 제공하여 초기에 발생하는 인지 노이즈를 최소화합니다. 두 번째 단계에서는 '선호되는 교사' 모델이 온-정책 증류를 통해 계획 능력을 전달하며, 명시적인 기호 상태를 사용하여 학생 모델의 시각적 탐색 과정을 지도합니다. 중요한 점은 기호 데이터는 학습 과정에서만 사용되며, 추론 과정은 순전히 시각 정보만을 활용하도록 설계되었습니다. 시각 계획 벤치마크 실험 결과, MGSD는 4B 및 8B 모델 모두에서 시각 계획 성능을 꾸준히 향상시켰습니다. 각각 19.3%와 18.4%의 매크로 평균 증가율을 보였습니다. 결과적으로, 개발된 모델은 기호 입력 기반 모델의 상한선에 더 가까워졌으며, 추가적인 분석과 진단 실험을 통해 성능 향상이 시각적 상태 복구 및 최적 경로 추론 모두에서 비롯되었음을 확인했습니다. 이러한 결과는 모달리티 격차를 인식하는 자기 증류가 모델이 실행 가능한 상태를 인지하는 방식뿐만 아니라, 추론된 구조에 대한 계획 과정을 개선한다는 것을 시사합니다. 코드: https://github.com/Oranger-l/MGSD

Original Abstract

While Vision-Language Models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this limitation to a perception--reasoning modality gap. Visual planning requires models to infer latent state structures from pixels and then reason over the recovered structure to produce valid actions, whereas symbolic planning directly leverages explicit representation. This discrepancy introduces two sequential bottlenecks: visual state recovery at the perception stage and multi-step planning at the reasoning stage. To address this, we propose MGSD, a two-stage modality-gap-aware self-distillation framework. First, a cold-start grounding stage establishes reliable visual state recovery before on-policy training. Second, a symbol-guided on-policy self-distillation stage transfers the privileged teacher's planning behavior to the student through token-level supervision on student-generated prefixes. Crucially, symbolic information is used only during training, while inference relies exclusively on visual inputs. Experiments on visual planning benchmarks show that MGSD consistently improves performance across different model scales, raising the macro average by 19.3% and 18.4%, respectively. The resulting models substantially reduce the gap to the upper bounds obtained with symbolic inputs. Ablation studies and diagnostic analyses further confirm that the gains arise from improvements in both visual state recovery and optimal-path reasoning. These results demonstrate that MGSD strengthens not only the recovery of actionable states from visual observations but also the ability to plan over the inferred structures. Code is available at https://github.com/Oranger-l/MGSD.

0 Citations
0 Influential
41.951858789481 Altmetric
0.0 Score
Original PDF
17

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!