2601.21306v1 Jan 29, 2026 cs.LG

모델 기반 강화 학습에서 탐색의 예상치 못한 어려움

The Surprising Difficulty of Search in Model-Based Reinforcement Learning

Wei-Di Chang
Wei-Di Chang
Citations: 346
h-index: 7
Mikael Henaff
Mikael Henaff
Citations: 676
h-index: 4
Brandon Amos
Brandon Amos
Citations: 123
h-index: 6
Gregory Dudek
Gregory Dudek
Citations: 69
h-index: 5
Scott Fujimoto
Scott Fujimoto
Citations: 59
h-index: 4

본 논문은 모델 기반 강화 학습(RL)에서의 탐색 문제를 연구합니다. 기존의 통념은 장기 예측과 누적 오차가 모델 기반 RL의 주요 장애물이라고 주장합니다. 우리는 이러한 관점에 도전하며, 탐색이 학습된 정책을 대체하는 단순한 방법이 아님을 보여줍니다. 놀랍게도, 모델의 정확도가 매우 높더라도 탐색이 성능을 저하시킬 수 있다는 것을 발견했습니다. 대신, 모델 또는 가치 함수의 정확도를 향상시키는 것보다 분포 변화를 완화하는 것이 더 중요하다는 것을 보여줍니다. 이러한 통찰력을 바탕으로, 효과적인 탐색을 가능하게 하는 핵심 기술을 제시하고, 여러 인기 있는 벤치마크 환경에서 최첨단 성능을 달성했습니다.

Original Abstract

This paper investigates search in model-based reinforcement learning (RL). Conventional wisdom holds that long-term predictions and compounding errors are the primary obstacles for model-based RL. We challenge this view, showing that search is not a plug-and-play replacement for a learned policy. Surprisingly, we find that search can harm performance even when the model is highly accurate. Instead, we show that mitigating distribution shift matters more than improving model or value function accuracy. Building on this insight, we identify key techniques for enabling effective search, achieving state-of-the-art performance across multiple popular benchmark domains.

4 Citations
1 Influential
3.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!