2606.24842v1 Jun 23, 2026 cs.AI

세분화된 월드 모델: 일반 에이전트를 위한 구조적 인증

World Models in Pieces: Structural Certification for General Agents

Xinyu Lu
Xinyu Lu
iscas.ac.cn
Citations: 185
h-index: 9
Tongxin Li
Tongxin Li
Citations: 11
h-index: 2
Yi Lu
Yi Lu
Citations: 52
h-index: 3
Yifei Wu
Yifei Wu
Citations: 4
h-index: 1

복잡한 환경에서, 에이전트는 모든 기능을 수행할 수 없으며, 불가피하게 월드 모델을 부분적으로 사용하여 특정한 능력에 집중하게 됩니다. 따라서, 기존의 균일적인 보장 방식은 중요한 문제점과 무관한 실패를 구별하는 데 한계가 있습니다. 우리는 이 한계를 공식화하고, 일반 에이전트가 범용적이지 않다는 것을 증명하여, 기존의 최악의 경우 분석 방법으로는 유용한 정보를 얻을 수 없음을 보여줍니다. 이를 극복하기 위해, 우리는 구조적 인증이라는 새로운 프레임워크를 제시합니다. 이는 특정 상태에서의 목표 달성 성능을 기반으로 에이전트의 내부 월드 모델에 대한 보장을 제공하며, 각 요소별로 검증합니다. 우리의 주요 기여는 실질적인 알고리즘입니다. 우리는 딥 컴포지셔널 목표를 사용하여 특정 상태를 필터링하는 알고리즘을 제시하고, 이러한 목표에 대해 일반 에이전트가 구조화된 월드 모델을 가지며, $\mathcal{O}(1/n) + \mathcal{O}(δ)$의 오차 경계를 가진다는 것을 증명합니다. 반대로, 이 경계는 작은 $δ$ 값에서 타이트하며, 이는 우리의 인증 방식에 의해 명시적으로 보장됩니다. 이러한 결과는 일반 에이전트의 신뢰성 있는 배포를 가능하게 하며, 장기적인 계획 수립이 가능한 특정 상태로 범위를 좁혀줍니다.

Original Abstract

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of critical bottlenecks and irrelevant failures. We first formalize this limitation by proving that general agents are not universal, rendering standard worst-case analysis uninformative. To overcome this, we introduce structural certification, a transition-local framework that maps bounded goal-conditioned performance to entry-wise guarantees on the agent's internal world model. Our main contribution is constructive. We provide algorithms that filter specific transitions using deep compositional goals and prove that a general agent on these goals has a structural world model with a $\mathcal{O}(1/n) + \mathcal{O}(δ)$ error bound. Conversely, this bound is tight in the small-$δ$ regime, whose existence is explicitly guaranteed by our certification. These results enable the certifiable deployment of general agents by localizing the specific transitions where long-horizon planning is reliable.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!