2606.27136v1 Jun 25, 2026 cs.AI

대규모 언어 모델 에이전트를 위한 경험 기반 규칙 및 정책의 통합 학습

Joint Learning of Experiential Rules and Policies for Large Language Model Agents

Chao Yu
Chao Yu
Citations: 8
h-index: 2
Shicheng Ye
Shicheng Ye
Citations: 4
h-index: 1

다단계 상호 작용 환경에서 LLM 에이전트는 축적된 상호 작용 경험을 효과적으로 활용하는 것이 중요한 과제입니다. 기존 연구에서는 이러한 경험을 일반적으로 두 가지 방식으로 사용합니다. 첫째, 자연어 규칙으로 모델 외부 저장하여 나중에 프롬프트에 활용하고, 둘째는 트래젝토리와 피드백을 사용하여 모델 파라미터를 업데이트합니다. 전자는 해석하기 쉽지만, 진화하는 정책과 동기화되지 않을 수 있습니다. 후자는 정책 전체를 개선하지만, 희소 보상 환경에서 발생하는 국지적인 오류에 대한 제한적인 수정만 제공합니다. 본 연구에서는 LLM 에이전트를 위한 경험 기반 규칙 및 정책의 통합 학습(JERP) 방법을 제안합니다. JERP는 동일한 상호 작용 트래젝토리를 사용하여 장기적인 경험 기반 규칙 풀과 정책을 모두 업데이트합니다. 의사 결정 시, JERP는 관련 규칙을 검색하고, 이를 에이전트에게 제공하며, 동시에 상호 작용 기록도 함께 활용합니다. 각 에피소드 이후에는 수집된 트래젝토리를 사용하여 정책을 최적화하고, 현재 실행 결과와 참조 성공 트래젝토리 간의 비교를 통해 규칙 풀을 수정합니다. 이러한 결합 방식은 규칙 풀이 진화하는 정책과 동기화되도록 유지하면서, 안정적이고 효과적인 행동을 점진적으로 모델 자체에 흡수하도록 합니다. AlfWorld 및 WebShop에서의 실험 결과는 JERP가 복잡한 상호 작용 작업에서 의사 결정 성능을 일관되게 향상시킨다는 것을 보여줍니다.

Original Abstract

For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model as natural-language rules for later prompting, or using trajectories and feedback to update the model parameters. The former is easy to interpret but can fall out of sync with the evolving policy; the latter improves the policy more broadly but provides only limited correction for local mistakes in sparse-reward settings. We present Joint Learning of Experiential Rules and Policies for LLM Agents (JERP), which updates a long-term experiential-rule pool and the policy from the same interaction trajectories. At decision time, JERP retrieves task-relevant rules and conditions the agent on them together with the interaction history. After each episode, it uses the collected trajectories both to optimize the policy and to revise the rule pool by comparing current rollouts with reference successful trajectories. This coupling keeps the rule pool aligned with the evolving policy while allowing stable and effective behaviors to be gradually absorbed into the model itself. Experiments on AlfWorld and WebShop show that JERP yields consistent gains in decision performance for complex interactive tasks.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!