시각적 행동 결과 추론 정렬을 통한 물리적 추론과 작업 일반화 능력 향상
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment
비전-언어 모델(VLMs)은 특히 새로운 작업 및 환경에서 발생하는 상호 작용적인 물리적 추론 상황에서 일반화 능력이 부족한 경향이 있습니다. 두 가지 주요 문제점이 두드러지는데, 첫째는 실제 물리 법칙과 모순되는 환각적 연쇄 사고(Chain-of-Thought, CoT) 추론이고, 둘째는 모델의 추론과 행동 간의 불일치입니다. 본 연구에서는 이러한 문제를 직접적으로 해결하기 위해 VAORA (Visual Action Outcome Reasoning Alignment), 즉 시각적 행동 결과 추론 정렬이라는 새로운 보상 체계를 제안합니다. VAORA는 두 가지 상호 보완적인 보상을 도입하는데, 첫째는 에이전트의 행동과 독립적으로 VLM의 추론을 시각적 맥락에 고정시키는 시각 정렬 보상(Visual Alignment Reward)이고, 둘째는 모델의 행동으로 인해 발생하는 시각적 결과에 기반하여 추론을 묶는 시각-행동 정렬 보상(Visual-Action Alignment Reward)입니다. 이러한 보상은 환각적인 CoT를 억제하고 추론과 행동 간의 격차를 줄이는 데 기여합니다. 또한, 사전 학습된 특정 환경 전문가 에이전트를 사용하여 성공 확률을 추정함으로써 부드럽고 밀도가 높은 보상을 적용하여 학습 안정성을 향상시켰습니다. PHYRE 및 Virtual Tool에서의 실험 결과는 VAORA가 새로운 작업 및 환경 설정에서 우수한 성능을 보이는 것을 확인시켜 주며, 이는 VAORA를 통해 기반 지식과 일반화 능력을 갖춘 물리적 인지 능력을 유도할 수 있음을 보여줍니다.
Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes are prominent: hallucinated chain-of-thought (CoT) reasoning that contradicts physical reality, and misalignment between the model's reasoning and actions. We present VAORA (Visual Action Outcome Reasoning Alignment), a novel reward design that directly addresses both issues. VAORA introduces two complementary rewards: Visual Alignment Reward, which anchors VLM reasoning to the visual context independent of the agent action itself, and Visual-Action Alignment Reward, which grounds reasoning in the visual outcome induced by the model's action. Together, these rewards suppress hallucinated CoT and reduce the gap between reasoning and behavior. To improve training stability, we further employ smooth, dense rewards by estimating success probabilities using a pre-trained in-domain expert agent. Experiments on PHYRE and Virtual Tool support our performances across novel-task and unseen-environment settings, confirming that grounded and generalizable physical intelligence can be induced through VAORA.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.