언어와 기호 표현 간 모달리티 전환을 통한 공간 추론
Spatial Reasoning via Modality Switching Between Language and Symbolic Representation
인간의 추론은 본질적으로 다중 양상을 지닙니다. 문제가 어려워지면, 우리는 종종 단어만으로 생각하지 않고, 대신 그림이나 격자 등을 활용하여 근본적인 개념적 구조를 이해하고 오류를 방지합니다. 이러한 전제하에, 본 연구는 다음 두 가지 질문을 탐구합니다: (a) 다단계 텍스트-공간 이야기를 기하학적 특징을 고려한 모달리티(예: 레이아웃 또는 격자)로 변환하는 것이 자연어 기반 추론보다 추론 능력을 향상시키는지 여부; 그리고 (b) 모델이 언제 자연어 추론에 의존해야 하고, 언제 구조화된 모달리티로 전환해야 하는지 판단할 수 있는지 여부. 우리는 신뢰도 및 복잡성 신호를 기반으로 한 전환 메트릭을 도입하여, 공간 이야기를 구조로 변환하는 것이 성능 향상에 얼마나 기여할 가능성이 높은지를 추정합니다. 이는 대규모 언어 모델(LLM)의 추론 과정에서 체계적인 모달리티 선택을 위한 첫걸음입니다. 실험 결과, 자연어 기반 추론에서 격자 기반 표현으로 전환하면 LLM의 성능이 최대 42%까지 향상되는 것으로 나타났으며, 이는 추론 결과를 결정하는 데 있어 모달리티 선택의 중요성을 강조합니다.
Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasoning by sketching diagrams or drawing grids to understand the underlying conceptual structure and avoid mistakes. Building on this premise, our research investigates: (a) whether grounding multi-hop textual-spatial stories into geometry-aware modalities, such as layouts or grids, improves reasoning compared to natural language-based inference; and (b) whether a model can decide when to rely on natural language reasoning and when to switch to a structured modality. We address these questions by introducing a switching metric based on trustworthiness and complexity signals, which estimates when grounding a spatial story into structure is likely to improve performance. This takes a first step toward principled modality selection in Large Language Model (LLM) reasoning. Across our settings, switching from natural language-based reasoning to a grid-based representation improves LLM performance by up to 42\%, highlighting the importance of modality choice in shaping reasoning outcomes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.