2604.20601v1 Apr 22, 2026 cs.AI

목표 조건 강화 학습을 이용한 지시 따르기 작업에서의 자가 학습 계획 추출

Self-Guided Plan Extraction for Instruction-Following Tasks with Goal-Conditional Reinforcement Learning

Alexey Skrynnik
Alexey Skrynnik
Citations: 451
h-index: 13
Z. Volovikova
Z. Volovikova
Citations: 77
h-index: 5
N. Sorokin
N. Sorokin
Citations: 46
h-index: 4
Dmitriy Lukashevskiy
Dmitriy Lukashevskiy
Citations: 0
h-index: 0
A. Panov
A. Panov
Citations: 1
h-index: 1

본 논문에서는 지시 따르기 작업을 위한 프레임워크인 SuperIgor을 소개합니다. 기존 방법들이 미리 정의된 하위 작업에 의존하는 것과 달리, SuperIgor은 언어 모델이 자체 학습 메커니즘을 통해 고수준 계획을 생성하고 개선할 수 있도록 하여, 수동 데이터셋 주석의 필요성을 줄입니다. 저희의 접근 방식은 반복적인 공동 훈련을 포함합니다. 강화 학습(RL) 에이전트는 생성된 계획을 따르도록 훈련되고, 언어 모델은 RL 피드백과 선호도를 기반으로 이러한 계획을 조정하고 수정합니다. 이를 통해 에이전트와 계획 모두가 함께 개선되는 피드백 루프가 형성됩니다. 저희는 복잡한 역학 및 확률적 특성을 가진 환경에서 저희의 프레임워크를 검증했습니다. 결과는 SuperIgor 에이전트가 기준 방법보다 지시를 더 엄격하게 준수하며, 동시에 이전에 보지 못한 지시에 대한 강력한 일반화 능력을 보여준다는 것을 나타냅니다.

Original Abstract

We introduce SuperIgor, a framework for instruction-following tasks. Unlike prior methods that rely on predefined subtasks, SuperIgor enables a language model to generate and refine high-level plans through a self-learning mechanism, reducing the need for manual dataset annotation. Our approach involves iterative co-training: an RL agent is trained to follow the generated plans, while the language model adapts and modifies these plans based on RL feedback and preferences. This creates a feedback loop where both the agent and the planner improve jointly. We validate our framework in environments with rich dynamics and stochasticity. Results show that SuperIgor agents adhere to instructions more strictly than baseline methods, while also demonstrating strong generalization to previously unseen instructions.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!