2607.01225v1 Jul 01, 2026 cs.LG

최적 이하 데모로부터의 언어 기반 비판 학습

Language-Critique Imitation Learning from Suboptimal Demonstrations

Kenneth Marino
Kenneth Marino
Citations: 82
h-index: 4
Dai-Jie Wu
Dai-Jie Wu
Citations: 52
h-index: 3
Chih-Han Yang
Chih-Han Yang
Citations: 12
h-index: 3
Yunao Huang
Yunao Huang
Citations: 40
h-index: 1
Ping-Chun Hsieh
Ping-Chun Hsieh
Citations: 45
h-index: 4
Shao-Hua Sun
Shao-Hua Sun
Citations: 51
h-index: 4

최적 이하 데모를 사용한 모방 학습 연구는 일반적으로 신뢰도 추정치, 판별기 점수 또는 중요도 가중치와 같은 압축된 감독 신호에 의존합니다. 이러한 스칼라 신호는 본질적으로 제한적이며, 작업 진행 상황, 실패 모드 또는 수정 조치에 대한 명시적인 중간 추론을 표현할 수 없습니다. 우리는 최적 이하 데모로부터의 모방 학습을 위한 언어 기반 비판 프레임워크를 제안합니다. 이 프레임워크는 자연어를 구조화된 감독 신호로 활용하여, 풍부한 피드백이 스칼라 값으로 축소되는 현상을 방지합니다. 우리의 방법은 먼저 현재 진행 상황을 명시적으로 설명하고, 최적 이하 행동을 식별하며, 세밀한 수정 지침을 제공하는 데모로부터 언어 레이블을 구성합니다. 그런 다음, 이러한 구조화된 신호를 스칼라 값으로 축소하지 않고 정책을 직접 훈련시키는 언어 비판 손실 함수를 도입하고, 이를 행동 복제 및 확산 정책에 적용하여 LC-BC 및 LC-DP를 구현합니다. 또한, 제안된 목적 함수가 표준 가정 하에서 전문가 성능 격차의 상한을 제공한다는 이론적 결과를 제시합니다. 경험적으로, 우리는 내비게이션, 조작 및 게임플레이를 포함하는 다양한 연속 제어 작업에서 우리의 방법을 평가했으며, 그 결과 우리의 방법은 강력한 모방 학습 및 오프라인 강화 학습 기준 모델보다 일관되게 우수한 성능을 보였습니다. 이러한 결과는 자연어가 최적 이하 데이터로부터 견고한 정책을 학습시키는 데 강력하고 구조화된 형태의 감독 신호로 사용될 수 있음을 보여줍니다.

Original Abstract

Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights. These scalar signals are inherently limited, as they cannot explicitly express intermediate reasoning about task progress, failure modes, or corrective actions. We propose a language-critique framework for imitation learning from suboptimal demonstrations that instead leverages natural language as a structured supervision signal, avoiding the collapse of expressive feedback into scalars. Our method first constructs language labels from demonstrations that explicitly describe current progress, identify suboptimal behaviors, and provide fine-grained corrective guidance. We then introduce a language-critique loss that directly trains policies using these structured signals without reducing them to scalars, and instantiate it for both behavior cloning and diffusion policies, yielding LC-BC and LC-DP. We further provide a theoretical result showing that the proposed objective upper-bounds the expert performance gap under standard assumptions. Empirically, we evaluate on diverse continuous control tasks spanning navigation, manipulation, and gameplay, where our methods consistently outperform strong imitation learning and offline reinforcement learning baselines. These results demonstrate that language can serve as a powerful and structured form of supervision for learning robust policies from suboptimal data.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!