2606.24669v1 Jun 23, 2026 cs.AI

LaGO: 온라인 강화 학습을 위한 잠재적 행동 가이드

LaGO: Latent Action Guidance for Online Reinforcement Learning

Tianhe Wu
Tianhe Wu
Citations: 128
h-index: 3
Kuan Liu
Kuan Liu
Citations: 0
h-index: 0
Rencong Huang
Rencong Huang
Citations: 70
h-index: 1

대규모 언어 모델(LLM)은 계획 및 순차적 의사 결정에 강력한 잠재력을 보여주었지만, 기존 연구에서는 LLM을 직접 제어기로 사용하는 경우가 많습니다. 이는 정확한 행동 생성이 필요하며 실제 환경에서 신뢰성이 떨어질 수 있습니다. 본 논문에서는 온라인 강화 학습을 위한 잠재적 행동 가이드(LaGO) 프레임워크를 제시합니다. LaGO는 사전 훈련된 LLM을 명시적인 계획기 또는 제어기가 아닌, 온라인 정책 최적화를 부드럽게 안내하는 잠재적 행동의 선행 정보로 활용합니다. 이산 제어를 위한 표준 벤치마크인 CLEVR-Robot과 연속 제어를 위한 벤치마크인 Meta-World에서 실험한 결과, LaGO는 Vanilla PPO보다 보상과 성공률 모두에서 일관되게 향상된 성능을 보였습니다. 특히, LaGO는 CLEVR-Robot에서 평균 성공률을 15.1%에서 27.2%로, Meta-World에서 2.7%에서 15.2%로 증가시켰습니다. 또한 분석 결과, 더 강력한 사전 훈련된 LLM이 더욱 효과적인 가이드 역할을 한다는 것을 확인했으며, 이는 LLM 지식이 계획 및 온라인 의사 결정을 향상시키는 데 기여할 수 있음을 시사합니다.

Original Abstract

Large language models (LLMs) have shown strong potential for planning and sequential decision-making, but prior work often relies on using them as direct controllers, which requires precise action generation and can be unreliable in practice. This paper proposes Latent Action Guidance for Online Reinforcement Learning (LaGO), a framework that uses a pretrained LLM as a latent action prior to softly guide online policy optimization, rather than treating the LLM as an explicit planner or controller. Experiments on both a discrete-control benchmark, CLEVR-Robot, and a continuous-control benchmark, Meta-World, demonstrate that LaGO consistently improves both reward and success rate over Vanilla PPO. In particular, LaGO increases the average success rate from 15.1% to 27.2% on CLEVR-Robot and from 2.7% to 15.2% on Meta-World. Our analysis further shows that stronger pretrained LLMs provide more effective guidance, suggesting that LLM knowledge can improve planning and online decision-making.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!