2606.19990v1 Jun 18, 2026 cs.AI

보상: 구체화된 세계 모델을 위한 에이전트

Reward as An Agent for Embodied World Models

Yongxuan Lv
Yongxuan Lv
Citations: 21
h-index: 2
Fei Wang
Fei Wang
Citations: 2
h-index: 1
Shan You
Shan You
Citations: 2,298
h-index: 19
Pu Li
Pu Li
Citations: 0
h-index: 0
Zhigang Lin
Zhigang Lin
Citations: 0
h-index: 0
Qiang Wu
Qiang Wu
Citations: 0
h-index: 0

강화 학습(RL)은 세계 모델을 개선하는 데 유망한 도구로 자리 잡았지만, 기존 방법들은 대부분 훈련 데이터 분포 근처의 보수적인 시뮬레이션을 기반으로 하여 탐색, 행동 다양성 및 풍부한 동적 발견 능력을 제한합니다. 본 연구에서는 이러한 보수적인 패러다임을 비판적으로 검토합니다. 우리는 핵심적인 제약이 탐색 자체가 아니라, 광범위한 탐색을 지원하기 위한 신뢰할 수 있는 검증 전략의 부족이라고 주장합니다. 신뢰할 수 있는 검증 없이는 확장된 탐색은 정책이 실제 개선 없이 불완전한 보상을 악용하는 '보상 해킹'에 매우 취약해집니다. 이러한 동기를 평가하기 위해, 본 연구에서는 물리적 타당성과 작업 완료가 복잡한 역학 하에서 확장 가능한 강화 학습을 위한 엄격한 테스트 환경을 제공하는 구체화된 세계 모델에 우리의 방법을 적용했습니다. 검증 측면에서, 우리는 생성된 행동을 능동적으로 평가하여 강력한 보상 신호를 제공하고 분포 변화(distribution shifts) 하에서의 보상 해킹을 완화하는 에이전트 기반의 보상 프레임워크인 '보상: 에이전트'를 소개합니다. 탐색 측면에서, 우리는 DynDiff-GRPO라는 동적 인지 롤아웃 다양화를 통해 명시적으로 행동 공간 탐색을 확장하여 경로를 다양화하고, 상태-행동 범위를 넓히며, 보수적인 롤아웃 방식으로는 얻기 어려운 풍부한 구체화된 행동을 유도합니다. '보상: 에이전트'와 DynDiff-GRPO를 통합함으로써, 우리는 훨씬 더 다양화된 샘플링을 통해 더욱 신뢰할 수 있는 기반 위에서 강화 학습을 가능하게 하며, 보상 해킹을 효과적으로 완화하면서 여러 공개 세계 모델에 걸쳐 상당한 정확도 향상을 달성합니다. 이를 통해 광범위한 탐색이 강력한 검증에 기반하여 성공적으로 확장될 수 있음을 입증합니다.

Original Abstract

While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training distribution, limiting exploration, behavioral diversity, and richer dynamic discovery. In this work, we challenge this conservative paradigm. We argue that the core limitation is not exploration itself, but the lack of reliable verification strategies to support broader exploration. Without reliable verification, expanded exploration becomes highly susceptible to reward hacking, where policies exploit imperfect rewards without achieving genuine improvement. To evaluate this motivation, we instantiate our method in embodied world models, where physical plausibility, and task completion provide a rigorous testbed for scalable RL under complex dynamics. On the verification side, we introduce Reward as an Agent, an agentic reward framework that actively evaluates generated behaviors to provide robust reward signals and mitigate reward hacking under distribution shifts. On the exploration side, we introduce Dynamic-Aware Rollout Diversification through DynDiff-GRPO, which explicitly expands action-space exploration to diversify trajectories, broaden state-action coverage, and encourage richer embodied behaviors beyond conservative rollout regimes. By unifying Reward as an Agent with DynDiff-GRPO, we enable RL on a more reliable reward foundation with substantially diversified sampling, effectively mitigating reward hacking while yielding significant accuracy gains across multiple open-source world models, thereby demonstrating that broader exploration can scale successfully when grounded in robust verification.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!