SKILLC: 대조 학습 기반 신뢰 할당을 통한 LLM 에이전트의 자율적인 기술 내재화 학습
SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment
구조화된 기술 프롬프트는 장기적 시야를 가진 에이전트 강화 학습(RL)에서의 탐색 능력을 향상시킵니다. 기술 기반 강화 학습 방법은 추론 단계에서 외부 기술을 유지하는 반면, 기술 내재화 강화 학습 방법은 훈련 과정에서 이를 제거하여 자율적인 성능을 가능하게 합니다. 그러나 기존의 내재화 방법은 단순히 커리큘럼 제어를 위해 기술 유용성을 비교하는 데만 사용되어 정책 업데이트를 변경하지 못하며, 기술 의존적 성공과 자율적인 성공을 구별할 수 없습니다. 본 연구에서는 대조 기반 기술 신뢰 할당(CSCA)에 기반한 프레임워크인 SkillC를 제안합니다. SkillC는 동일 정책 업데이트 내에서 활성 기술 유형의 작업에 대해 쌍으로 구성된 기술 적용 및 비적용 롤아웃을 샘플링하고, 이들의 작업 수준 대비를 사용하여 글로벌 순위를 유지하면서 기술 비적용 성공 방향으로 일방적인 수정을 가하는 이중 스트림 어드밴티지 추정기를 통해 최적화 과정에 직접적인 학습 신호를 제공합니다. 더욱이 검증 레벨의 신호를 활용하여 속성 강도, 롤아웃 할당 및 단조적인 활성 집합 가지치기에 대한 적응형 커리큘럼을 구현합니다. ALFWorld와 WebShop에서의 실험 결과, SkillC는 실행 시간에 기술 접근 권한 없이 기존의 최고 성능을 보이는 기술 내재화 강화 학습 모델보다 각각 5.5% 및 4.4% 더 높은 성능을 보이며, 동시에 기술 기반 강화 학습 방법과 경쟁력 있는 수준을 유지합니다.
Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at inference, while skill-internalization RL methods withdraw them during training to enable autonomous performance. However, existing internalization approaches only use skill-helpfulness contrast for curriculum control, leaving the policy update unchanged and unable to distinguish skill-dependent from autonomous success. We propose SkillC, a framework based on Contrastive Skill Credit Assignment (CSCA) that converts this contrast into a direct learning signal for internalization. \textsc{SkillC} samples paired skill-injected and skill-free rollouts for tasks from active skill types within the same policy update, and injects their task-level contrast into optimization via a dual-stream advantage estimator that preserves global ranking while applying a one-sided correction toward skill-free success. A smoothed validation-level signal further drives an adaptive curriculum over attribution strength, rollout allocation, and monotonic active-set pruning. Experiments on ALFWorld and WebShop show that, without runtime skill access, SkillC surpasses the strongest prior skill-internalization RL baseline by 5.5\% and 4.4\%, respectively, while remaining competitive with skill-augmented RL methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.