2605.26403v1 May 26, 2026 cs.AI

정적 환경에서 교정된 상호작용 강화 학습으로: 정렬된 시뮬레이터를 활용한 다중 회전 대화에서의 분포 변화 완화

From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator

Muzhao Tian
Muzhao Tian
Citations: 44
h-index: 5
Zisu Huang
Zisu Huang
Citations: 71
h-index: 6
Xiaohua Wang
Xiaohua Wang
Citations: 425
h-index: 9
Changze Lv
Changze Lv
Fudan Univerisity
Citations: 649
h-index: 11
Xiaoqing Zheng
Xiaoqing Zheng
Citations: 624
h-index: 10
Jiakang Yuan
Jiakang Yuan
Fudan University
Citations: 570
h-index: 14
Kaitao Song
Kaitao Song
Citations: 23
h-index: 3
Tao Chen
Tao Chen
Citations: 38
h-index: 3

연구 커뮤니티의 오랜 목표는 고도의 상호작용성을 갖춘 LLM 기반 대화 에이전트를 개발하는 것입니다. 최근 연구에서는 고정된 오프라인 로그(정적 환경 강화 학습)를 기반으로 정책을 최적화하거나, 프롬프트 기반 시뮬레이터를 사용하는 방법(상호작용 강화 학습)에 초점을 맞추고 있습니다. 본 논문에서는 이론적으로 두 가지 패러다임 모두가 근본적인 한계를 가지고 있으며, 이는 대화 기록의 분포 변화(training 과정에서 관찰된 대화 기록과 실제 대화에서 발생하는 기록 간 불일치) 때문입니다. 이러한 변화는 회전 수에 따라 제곱으로 증가하며, 대화 품질을 심각하게 저하시킵니다. 구체적으로, 이 변화는 두 가지 주요 원인으로부터 발생합니다: (i) 정책 유도 변화(policy-induced shift), 이는 정적 기록이 아닌 self-generated trajectory를 기반으로 학습하기 때문에 발생하는 현상입니다; 그리고 (ii) 시뮬레이터 유도 변화(simulator-induced shift), 이는 시뮬레이션된 행동과 실제 인간의 행동 간의 불일치에서 비롯됩니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 상호작용 강화 학습을 시뮬레이터 정렬과 결합한 통합 프레임워크인 Calibrated Interactive RL을 제안합니다. 우리의 접근 방식은 시뮬레이터를 인간과의 상호 작용 패턴에 맞춰 조정함으로써 시뮬레이션-실제 격차를 줄이고, 누적되는 분포 변화를 완화합니다. 여러 대화 작업에서의 실험 결과는 이론적 분석을 뒷받침하며, (i) 상호작용 강화 학습은 정책 분포 변화를 완화하여 정적 환경 기반의 baseline보다 훨씬 우수한 성능을 보이며, (ii) 제안하는 alignment 방법을 통해 시뮬레이터를 교정함으로써 시뮬레이션-실제 격차를 더욱 줄여, 최첨단 수준의 downstream 성능을 달성합니다.

Original Abstract

A long-standing goal of the research community is to develop highly interactive LLM-based dialogue agents. Recent research focuses on optimizing policies based on fixed offline logs (Static Context RL) or using a prompt-based simulator (Interactive RL). In this work, we theoretically show that both paradigms are fundamentally limited by context distribution shift--a mismatch between dialogue histories observed during training and those encountered in real conversations. This shift compounds quadratically over turns and severely degrades dialogue quality. Specifically, we attribute this shift to two distinct sources: (i) policy-induced shift, arising from training on static histories rather than self-generated trajectories; and (ii) simulator-induced shift, stemming from discrepancies between simulated and real human behaviors. To address these challenges, we propose Calibrated Interactive RL, a unified framework that couples interactive RL with simulator alignment. By aligning the simulator with human interaction patterns, our approach reduces the sim-to-real gap and mitigates compounding distribution shifts. Experiments across multiple dialogue tasks confirm our theoretical analysis: (i) Interactive RL significantly outperforms the Static Context baseline by mitigating policy distribution shift; and (ii) calibrating simulators with our alignment method further bridges the sim-to-real gap, yielding state-of-the-art downstream performance.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!