2607.29617v1 Jul 31, 2026 cs.LG

언제 온-폴리시 상호작용이 도움이 되는가? 가치 기반 모방 학습에서의 표현적 절충 관계

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

P. Amortila
P. Amortila
Citations: 297
h-index: 8
Dylan J. Foster
Dylan J. Foster
Citations: 648
h-index: 12
Audrey Huang
Audrey Huang
Citations: 274
h-index: 5
V. Cevher
V. Cevher
Citations: 16,461
h-index: 62
Luca Viano
Luca Viano
Citations: 220
h-index: 9
Antoine Moulin
Antoine Moulin
Citations: 22
h-index: 3

모방 학습(IL)은 로봇 공학부터 언어 모델 훈련에 이르기까지 다양한 분야에서 활용되는 기술로, 에이전트가 전문가의 행동을 데모를 통해 학습하도록 하는 방법입니다. Behavior Cloning (BC)과 같은 일반적인 접근 방식은 복합적인 오류와 성능 정체의 문제를 겪는 경향이 있으며, 특히 학습자가 전문가의 정책을 완벽하게 표현할 수 없을 때 이러한 문제가 더욱 심각해집니다 (예: 지식 증류). 두 가지 방법이 경험적으로 성능 향상에 기여하는 것으로 알려져 있습니다. 첫째는 학습자의 자체 경로를 따라 전문가에게 상호 작용하여 정보를 얻는 것이고, 둘째는 정책을 직접 생성하기보다는 가치 함수 추정을 활용하는 것입니다. 본 연구에서는 이러한 개선 사항의 본질과 그 잠재적으로 놀라운 상관 관계를 조사합니다. 주요 발견은 전문가와의 상호 작용이 학습자의 표현 능력에 대한 요구 사항을 완화한다는 것입니다. 즉, 학습자는 전문가의 정책 자체를 구현해야 하는 엄격한 조건을 피하고, 단순히 전문가의 가치 함수를 실현할 수 있는 모델만 있으면 됩니다. 구체적으로, 저희는 OVI라는 대화형 온-폴리시 모방 학습 알고리즘을 소개합니다. OVI는 학습자가 전문가의 가치 함수를 표현할 수 있고 선형 최대화 오라클에 접근할 수 있다면 통계적으로 효율적이며 계산적으로도 효율적입니다. 또한, 상호 작용이 반드시 필요하다는 것을 보여주는 부정적인 결과를 제시합니다. 즉, 전문가의 가치 함수 구현 능력 외에 더 강력한 가정 없이 어떤 오프라인 모방 학습 알고리즘도 전문가 정책 클래스의 복잡성에 따라 확장되어야 합니다. 이러한 발견은 실험적으로 입증되었습니다. OVI는 오프라인 정책 기반 (BC), 대화형 정책 기반 (DAgger) 및 오프라인 가치 기반 모방 학습 방법보다 성능이 우수하며, 특히 학습자 네트워크의 표현력이 전문가의 것보다 현저히 낮을 때 가장 큰 개선 효과를 보입니다.

Original Abstract

Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when the learner cannot perfectly represent the expert's policy (as is typical, e.g., in distillation). Two interventions are widely understood empirically to improve performance: querying the expert interactively along the learner's own trajectories, and using value function estimation en route to generating a policy rather than directly fitting the expert's full action distribution. We investigate the nature of these improvements and their potentially surprising interplay. Our main finding is that expert interaction relaxes the representational demands on the learner: one only needs a model capable of realizing the expert's value function, bypassing the (often stricter) requirement of realizing the expert's policy itself. Concretely, we introduce OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle. We complement this with a negative result showing that interaction is necessary. Namely, without stronger assumptions beyond expert-value realizability alone, any offline IL algorithm must scale with the complexity of the expert policy class. Our findings bear out empirically. OVI outperforms offline policy-based (BC), interactive policy-based (DAgger), and offline value-based IL methods, with the largest gains when the learner network is substantially less expressive than the expert's.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!