구두 강화 학습을 이용한 대규모 언어 모델 개인화에서의 선호도 적응 학습
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning
자연어 사용자 선호도는 LLM(Large Language Model) 개인화를 위한 해석 가능한 인터페이스를 제공합니다. 그러나, 보편적인 선호도 요약은 특정 하위 작업과 관련 없는 정보를 포함하는 경우가 많습니다. 따라서 전체 선호도 요약을 직접 제공하면 컨텍스트 용량을 낭비하고, 작업 간의 방해 요소를 유발할 수 있으며, 수동으로 작업별 선호도 관점을 설계하는 것은 확장하기 어렵습니다. 본 연구에서는 extit{작업별 선호도 적응}을 연구합니다. 즉, 보편적인 사용자 선호도 요약과 하위 작업을 기반으로, 의사 결정에 필요한 충분한 정보를 유지하면서 불필요한 컨텍스트를 제거하는 작업 조건화된 표현을 도출합니다. 이를 위해, 우리는 훈련이 필요 없는 메타 학습 프레임워크인 extsc{AlignXada}를 제안합니다. extsc{AlignXada}는 재사용 가능한 텍스트 개선 정책을 유도하여 보편적인 선호도 요약을 작업별로 적응시킵니다. 개선 정책은 구두 강화 학습을 통해 메타 러너에 의해 반복적으로 최적화됩니다. 13개의 작업과 세 가지 하위 모델(총 39개 작업-모델 조합)에 대해, extsc{AlignXada}는 평균 3.82점의 성능 향상을 달성했으며, 33개 셀에서 성능이 개선되고 원래 프로필 토큰의 22.8%만 유지되었습니다. 또한 RAG(Retrieval-Augmented Generation)를 능가하는 결과를 36개 셀에서 보여주었습니다. 추가적인 충실도 분석 결과, 개선된 프로필은 여전히 원본 선호도에 기반하고 있으며 작업 관련 개인화 신호를 보존한다는 것을 확인했습니다. 이는 프로필 측면의 적응이 평생 개인화 에이전트를 위한 보편적 메모리 구축을 위한 실용적인 보완 방법임을 시사합니다.
Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale. In this work, we study \emph{task-specific preference adaptation}: given a universal user preference summary and a downstream task, derive a task-conditioned representation that preserves sufficient decision-relevant evidence while removing redundant context. To this end, we propose \textsc{AlignXada}, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The refinement policy is iteratively optimized by a meta learner through verbal reinforcement learning. Across 13 tasks and three downstream models (39 task--model cells), \textsc{AlignXada} achieves an average gain of 3.82 points, improving 33 cells while retaining only 22.8\% of the original profile tokens and outperforming RAG in 36 cells. An extended faithfulness analysis further shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.