피드백 루프를 통한 학습: 언어 기반 강화 학습에서 경험 추출부터 인사이트 거버넌스까지
Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning
훈련 없이 작동하는 언어 기반 강화 학습은 LLM 에이전트가 동적 작업 결과, 시장 수익 또는 수요 예측과 같은 객관적인 세계 피드백으로부터 학습할 수 있도록 합니다. 이 방법은 경험에서 언어 규칙을 추출하여 컨텍스트로 주입함으로써 에이전트의 동작을 매개변수 변경 없이 업데이트합니다. 그러나 비정상 환경에서는 이러한 에이전트가 '보존-망각' 딜레마에 직면하는데, 이는 오래된 인사이트를 유지하면 부정적인 영향(negative transfer)을 초래하고, 이를 버리면 조건이 재발할 때 심각한 망각(catastrophic forgetting)을 야기합니다. 우리는 이 딜레마를 해결하기 위한 네 가지 요건 -- 결과 중심 평가, 지속적인 구조화된 증거, 비단조적 지식 라이프사이클, 그리고 합성 거버넌스 --을 제시하고, 기존 방법들이 경험 추출에 과도하게 투자하는 반면 인사이트 거버넌스는 상대적으로 소홀히 한다는 점을 보여줍니다. 우리는 규칙, 증거 및 기술이라는 세 계층 아키텍처를 제안하며, 이들은 피드백 기반 큐레이션 루프에 의해 연결되어 거버넌스 격차를 해소합니다. 규칙은 세계 결과에서 추출된 경험을 요약하고, 증거는 각 규칙의 신뢰성을 에피소드별로 추적하며, 기술은 어떤 규칙을 적용할지, 충돌을 어떻게 해결할지, 그리고 언제 회피할지를 결정합니다. 본 연구에서는 금융 예측이라는 사례를 통해 세계 피드백이 풍부하고 노이즈가 많으며 비정상적인 환경에서, 동일한 축적된 경험이 큐레이션 루프의 유무에 따라 기존 성능보다 저조하거나 정확도와 위험 조정 수익을 크게 향상시킬 수 있음을 보여줍니다.
Training-free verbal reinforcement learning enables LLM agents to learn from world feedback -- objective signals such as dynamic task outcomes, market returns, or demand forecasts -- by extracting verbal rules from experience and injecting them as context, updating the agent's behavior without parameter changes. However, in non-stationary environments these agents face a retention-forgetting dilemma: retaining stale insights causes negative transfer, while discarding them causes catastrophic forgetting when conditions recur. We identify four requirements for navigating this dilemma -- outcome-driven evaluation, persistent structured evidence, non-monotonic knowledge lifecycle, and compositional governance -- and show that existing methods invest heavily in experience extraction while underinvesting in insight governance. We propose a three-layer architecture -- rules, evidence, and skills -- connected by a feedback-driven curation loop that closes the governance gap. Rules capture distilled experience from world outcomes; evidence logs track each rule's reliability across episodes; skills govern which rules to apply, how to resolve conflicts, and when to abstain. On financial forecasting as a case study, where world feedback is naturally abundant, noisy, and non-stationary, we show that the same accumulated experience either degrades performance below the zero-shot baseline or dramatically improves accuracy and risk-adjusted returns, depending on whether the curation loop is present.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.