Recon: 재구성 기반 추론 합성 - 사용자 모델링을 위한 접근 방식
Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling
사용자 모델링은 언어 모델(LM)을 사용하여 과거의 문맥-행동 쌍 데이터(예: 대화 내용)를 바탕으로 개인의 행동을 모방하는 것을 목표로 하며, 이는 행동 과학, 인간-AI 협업 및 시장 조사와 같은 환경에서 사용자를 시뮬레이션할 수 있도록 합니다. 최근 연구에서는 이러한 데이터를 합성된 추론 과정을 추가하여 개선해왔으며, 일반적으로 문맥과 행동 모두를 조건으로 사용하여 생성합니다. 그러나 이러한 방식은 실제 추론이라기보다는 사후적 정당화에 해당하며, 생성된 추론 과정이 행동을 정당화하도록 보장하지만, 근본적인 잠재적인 인과 결정 경로를 반드시 반영하지는 않습니다. 본 연구에서는 'Recon'이라는 새로운 방법을 제안합니다. Recon은 행동 재구성을 사용하여 추론 과정을 평가하며, 예측력을 기준으로 품질을 측정합니다. 구체적으로, 주어진 문맥과 후보 추론 과정에 대해 재구성 모델이 행동을 예측하고, 재구성의 정확도가 추론 과정의 품질을 결정합니다. 네 가지 영역에서 실험한 결과, Recon은 표준적인 사후적 정당화 방법인 'Backward Synthesis'보다 54.7% 더 높은 성능을 보였습니다. 또한, Recon으로부터 얻은 보상을 사용하여 추론 합성 모델을 학습시키면 사용자 모델링 성능이 향상되며, 기준 모델 대비 최대 70.0%의 성능 향상을 달성할 수 있었습니다. 추가적으로, Recon으로 생성된 추론 과정이 다른 모델에서도 활용 가능하며, 재구성 모델 자체를 넘어선 사용자 모델링 성능 향상을 가져올 수 있음을 확인했습니다. 본 연구는 사후적 정당화만으로는 효과적인 추론 합성을 할 수 없으며, 유용하고 해석 가능한 추론은 문맥으로부터 자연스럽게 행동을 이끌어내야 함을 보여줍니다.
User modeling aims to use language models (LMs) to mimic an individual's behavior from a corpus of past context-action pairs (e.g., conversation turns), enabling the simulation of users in settings like behavioral science, human-AI collaboration, and market research. Recent approaches augment these corpora with synthesized reasoning traces, typically generated by conditioning on both context and action. However, such conditioning constitutes post-hoc rationalization rather than reasoning: the trace is guaranteed to justify the action, but may not encode the underlying latent causal decision paths. We propose Recon, which uses action reconstruction to score reasoning traces by their predictive power: given a context and candidate reasoning, a reconstruction model predicts the action, and reconstruction fidelity determines reasoning quality. Across four domains, Recon achieves a 54.7% win rate over Backward Synthesis, a standard post-hoc rationalization baseline. Further, we find that training a reasoning synthesis model with rewards derived from Recon improves downstream user modeling performance, achieving a win rate of up to 70.0% over baselines. We further show that Recon-synthesized reasoning transfers across models, and improves user modeling beyond the reconstruction model. Our work demonstrates that post-hoc rationalization is insufficient for reasoning synthesis, and that useful and interpretable reasoning should naturally elicit the action from the context.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.