정신 세계 모델링
Mental World Modeling
세계 모델은 계획 및 행동을 위한 예측 기반을 제공하지만, 기존의 방식들은 주로 물리적인 질문에 답하는 데 집중합니다: 무엇이 어디에 있는지, 그리고 그것이 어떻게 변화할 것인지. 그러나 인간의 행동은 숨겨진 정신 상태(개인이 믿는 것, 원하는 것, 의도하는 것, 느끼는 것, 사회적으로 용인되는 것으로 생각하는 것)에 의해 주도됩니다. 따라서 물리적 장면을 추적하지만 각 에이전트가 해당 장면 자체에 대해 무엇을 알고 있는지, 무엇을 믿고 있는지를 고려하지 않는 모델은 겉보기에는 적절해 보이는 상황에서도 잘못된 행동을 예측할 수 있습니다. 우리는 정신 변수를 세계 모델의 핵심 구성 요소로 간주하는 일반적인 이론적 프레임워크인 '정신 세계 모델링 (Mental World Modeling, MWM)'을 제안합니다. MWM은 물리적-정신적 세계 상태를 결합하고, 대상에 특화된 부분 관찰 결과를 제공하며, 후보 행동이 어떻게 양쪽 구성 요소를 동시에 업데이트하는지를 시뮬레이션합니다. 우리는 이 프레임워크를 MENTIS라는 학습이 필요 없고 완전히 검토 가능한 기본 모델로 구현했습니다. MENTIS는 프로세스를 상태 분석, 대상-관찰 생성, 행동 분해, 결합된 물리적 및 정신적 변환, 그리고 분기 수준의 가치 평가로 구성합니다. 텍스트, 이미지, 오디오-비디오 스토리로 구성된 수동으로 제작되고 품질 관리가 이루어진 상황 판단 시나리오 데이터 세트를 사용하여 8개의 최신 LLM 기반 세계 모델을 대상으로 실험한 결과, 명시적으로 정신 상태를 모델링하는 것이 인간의 의사 결정을 예측하는 데 필수적이라는 것을 확인했습니다. 더 깊이 있는 분석은 현재 정신 세계 모델링의 한계를 더욱 드러냅니다. 우리는 MWM이 세계 모델링의 다음 단계가 될 것으로 기대합니다. 즉, 물리적인 장면을 시뮬레이션하는 것에서 그 장면 안에서 행동하는 마음을 시뮬레이션하는 것으로 나아가는 것입니다.
World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.