교차 사고에서의 모달 격리 해소: 단계별 강화 학습을 통한 모달 전환 감독
Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement
교차 사고는 단일 다중 모드 모델이 텍스트 기반 추론과 시각적 생성 기능을 번갈아 사용하는 방식으로, 공간 및 물리적 작업에서 유망한 결과를 보여주었습니다. 그러나 복잡하고 긴 연쇄적인 시나리오에서, 우리는 근본적인 문제점인 '모달 격리' 현상을 발견했습니다. 이는 생성된 이미지가 텍스트 맥락과 어긋나고, 후속 텍스트가 시각적 증거를 무시하여 두 모드가 서로에게 진정으로 정보를 제공하지 않고 번갈아 가며 작동하는 현상입니다. 우리는 이 현상을 '모달 격리'라고 명명하고, 이는 모달 경계에서의 정보 손실 누적으로 인해 발생한다고 봅니다. 우리는 각 추론 주기를 기본 연산으로 분해하고, 모달 전환 시 발생하는 교차 모드 환각(텍스트-이미지) 및 시각적 활용 부족(이미지-텍스트)을 정량화하는 '모달 전환 손실'을 정의했습니다. 우리는 MoTiF (Modality Transition Fidelity)라는 두 단계의 학습 프레임워크를 제안합니다. 이 프레임워크는 모달 전환을 직접 최적화하며, 첫 번째 단계인 Reflective SFT는 모델이 오류가 있는 시각적 출력을 감지하고 복구하도록 훈련시키고, 두 번째 단계인 Flow-GRPO는 강화 학습을 통해 이미지 생성의 정확도를 향상시킵니다. MoTiF에서 사용되는 모든 학습 신호는 최종 작업의 정확도가 아닌 전환 수준에서의 충실도에 기반합니다. 네 가지 시각적 퍼즐 벤치마크를 통해, 이러한 전환 수준에서의 감독은 교차 모드 일관성과 최종 작업 정확도를 크게 향상시켰습니다. 이러한 결과는 효과적인 교차 사고가 단순한 확장 또는 최종 작업 최적화가 아닌, 모달 경계에서의 명시적인 구조적 감독이 필요하다는 것을 보여줍니다.
Interleaved thinking, where a unified multimodal model alternates between textual reasoning and visual generation, has shown promise on spatial and physical tasks. However, in complex long-chain scenarios, we identify a fundamental failure mode: generated images diverge from the textual context while subsequent text ignores the visual evidence, causing the two modalities to alternate without genuinely informing each other. We term this Modal Isolation and attribute it to compounding information loss at modality boundaries. We decompose each reasoning cycle into atomic operations and define modality transition loss, quantifying cross-modal hallucination (text-to-image) and visual utilization deficit (image-to-text) at each boundary. We propose MoTiF (Modality Tiransition Fidelity), a two-stage training framework that directly optimizes these transitions: Reflective SFT trains the model to detect and recover from erroneous visual outputs; Flow-GRPO improves image generation fidelity via reinforcement learning. All training signals in MoTiF derive from transition-level fidelity rather than end-task accuracy. Across four visual puzzle benchmarks, this transition-level supervision substantially improves both cross-modal coherence and final task accuracy. The results demonstrate that effective interleaved reasoning requires explicit structural supervision at modality boundaries, not merely scaling or end-task optimization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.