세계 모델 붕괴: 상전이로서의 현상
World-Model Collapse as a Phase Transition
물은 온도가 상승해도 외관상 변함없이 보이다가, 특정 임계점에서 갑자기 끓기 시작합니다. 본 연구에서는 장기적인 환경에서 작동하는 언어 기반 에이전트들이 내재적으로 사용하는 세계 모델에서도 유사한 상전이가 발생하는지 조사했습니다. 특정 파라미터 설정 하에서, 상태 로드 값을 미세하게 변경하거나 호라이즌 단계를 하나 추가하더라도 에이전트의 행동은 거의 변하지 않습니다. 하지만 임계 경계 근처에서는 동일한 작은 변화가 갑작스러운 세계 모델 붕괴를 초래합니다. 본 연구는 정확한 단계별 목표 상태를 가진 결정론적 작업 환경에서 이 효과를 분석했습니다. 상태 cardinality, 의존성 밀도, 호라이즌, 분기, 관찰 모드 및 변이율에 대한 광범위한 매개변수 탐색을 통해 상전이 다이어그램을 제시합니다: 해결된 안정 영역, 좁은 상전이 대역, 그리고 붕괴 지점. 단계별 추적 결과를 통해 세계 상태의 충실도가 행동 유효성보다 먼저 손상된다는 것을 확인했습니다. 즉, 에이전트는 단순히 잘못된 행동을 선택하는 것이 아니라, 왜곡된 세계 모델에 기반하여 행동하고 있습니다. 더 강력한 모델은 임계 경계를 이동시키지만, 질적인 상전이는 제거하지 못합니다. 본 연구 결과는 세계 모델 붕괴가 장기적인 환경에서 작동하는 에이전트의 성능 저하를 유발하는 중요한 병목 현상임을 시사합니다.
Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transition in their implicit world models. In some parameter settings, changing state load by a small amount, or adding a single step of horizon, leaves behavior nearly unchanged; near a critical boundary, the same small change causes a sudden world collapse. We study this effect in a deterministic task family with exact per-step gold state. A large grid search over state cardinality, dependency density, horizon, branching, observation mode, and mutation rate reveals a phase diagram: a solved plateau, a narrow transition band, and a collapse floor. Per-step traces show the mechanism: world-state fidelity fails before action validity, so the agent is not merely choosing a bad action; it is acting from a corrupted world. Stronger models translate the critical boundary but do not remove the qualitative transition. These results make world-model collapse a measurable bottleneck for long-horizon agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.