2606.05555v1 Jun 04, 2026 cs.LG

표현 학습은 확장 가능한 다중 작업 심층 강화 학습을 가능하게 한다

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

Johan Obando-Ceron
Johan Obando-Ceron
University of Montreal/ Mila
Citations: 752
h-index: 14
Pierre-Luc Bacon
Pierre-Luc Bacon
Citations: 3,433
h-index: 22
Scott Fujimoto
Scott Fujimoto
Citations: 59
h-index: 4
Aaron C. Courville
Aaron C. Courville
Citations: 123
h-index: 6
Pablo Samuel Castro
Pablo Samuel Castro
Citations: 48
h-index: 2
Lu Li
Lu Li
Citations: 25
h-index: 2

강화 학습(RL)을 다양한 다중 작업 환경으로 확장하는 것은 여전히 중요한 과제이다. 최근 모델 기반 RL의 발전은 뛰어난 성능을 달성했지만, 계획 및 복잡한 훈련 파이프라인에 의존하기 때문에 어떤 구성 요소가 확장성에 필수적인지 불분명하다. 본 연구는 이 질문을 재검토하고, 확장 가능한 다중 작업 RL의 주요 동인은 모델 기반 제어가 아니라 extit{표현 학습}이라고 주장한다. 특히, 예측적이고 모델 기반 표현과 고용량을 갖는 가치 함수 근사 방법을 결합하는 것만으로도 계획 없이도 강력한 성능을 달성할 수 있음을 보여준다. 본 연구에서는 간단한 모델-프리 알고리즘인 MR.Q를 보조적인 예측 목표와 함께 확장 가능한 액터-크리틱 아키텍처에 통합하여 평가하였다. 이 접근 방식은 최근의 세계 모델 기반 방법 및 다양한 심층 RL 기준 방법을 능가하며, 다중 작업 연속 제어 과제 전반에서 우수한 성능을 보인다. 또한, 계산 오버헤드를 크게 줄이고 실제 시간 효율성을 향상시킨다. 모델 용량을 늘리면 일관된 성능 향상이 관찰되었으며, 실험 분석 결과 예측적 표현 학습이 성능에 매우 중요하다는 것을 확인하였다.

Original Abstract

Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge. While recent advances in model-based RL achieve strong performance, they rely on planning and complex training pipelines, making it unclear which components are essential for scalability. We revisit this question and argue that the primary driver of scalable multitask RL is not model-based control, but \emph{representation learning}. In particular, we show that combining predictive, model-based representations with high-capacity value function approximation is sufficient to achieve strong performance, even without planning. We evaluate a simple model-free algorithm, MR.Q, coupled with auxiliary predictive objectives into a scalable actor-critic architecture. This approach outperforms a recent world-model-based method and a range of deep RL baselines across a diverse suite of multitask continuous control tasks, while significantly reducing computational overhead and improving wall-clock efficiency. We observe consistent improvements with increased model capacity and show through ablations that predictive representation learning is critical for performance.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!