PerMix-RLVR: 검증 가능한 보상 정렬을 통한 페르소나 표현력 보존
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
페르소나 프롬프팅은 특정 캐릭터를 부여하여 대규모 언어 모델(LLM)의 행동을 제어하고 지시 수행 능력을 향상시키는 데 널리 사용됩니다. 그러나 최적의 페르소나를 찾는 데 시간이 오래 걸리고, 그 영향이 출력 품질에 미치는 영향은 제대로 이해되지 않습니다. 기존 연구는 주로 추론 시간 전략을 통해 프롬프트 수준에서 이 문제를 해결하려고 시도했지만, 추가적인 계산 비용이 발생합니다. 본 연구에서는 추론 시간 프롬프트 검색을 피하기 위해, 학습 과정에서 페르소나 민감도를 해결하여, 다양한 페르소나에 적응하면서도 작업 성능을 유지하는 모델을 학습하는 것을 목표로 합니다. 특히, 검증 가능한 보상 기반 강화 학습(RLVR)이 페르소나 프롬프트에 대한 민감도를 체계적으로 감소시키지만, 결과 기반 최적화의 고유한 상충 관계를 드러낸다는 것을 발견했습니다. 즉, RLVR은 검증 가능한 목표를 가진 작업에서 견고성을 향상시키지만, 필요한 경우, 예를 들어 캐릭터 기반 역할극에서 페르소나 표현력을 저하시킬 수 있습니다. 이러한 한계를 극복하기 위해, 페르소나 견고성-충실도 간의 상충 관계를 완화하는 페르소나 혼합 RLVR(PerMix-RLVR) 전략을 제안합니다. PerMix-RLVR은 유해한 페르소나 변화에 대한 강력한 견고성을 유지하면서, 필요한 경우 충실한 페르소나 적용을 가능하게 합니다. 구체적으로, PerMix-RLVR은 MATH500 데이터셋에서 RLVR에 비해 페르소나 안정성 점수(PSS)를 +21.2% 향상시키고, PersonaGym 데이터셋에서 페르소나 충실도를 +11.4% 향상시킵니다.
Persona prompting has been widely adopted to steer large language models (LLMs) behavior and improve their instruction performance by assigning specific characters. However, identifying an optimal persona is time-consuming, and its impact on output quality remains poorly understood. Prior work has mainly addressed this issue at the prompt level via inference-time strategies, incurring additional computation. In this work, we avoid inference-time prompt search by tackling persona sensitivity during training, aiming to train models that adapt their behavior to diverse personas while preserving task performance. In particular, we find that reinforcement learning with verifiable rewards (RLVR) systematically reduces sensitivity to persona prompts, but also reveals an inherent trade-off of outcome-based optimization: while RLVR improves robustness on tasks with verifiable goals, it can also degrade persona expressivity when needed, e.g., in-character role-playing. To address this limitation, we propose PerMix-RLVR, a persona-mixed RLVR strategy that mitigates the persona robustness-fidelity trade-off, preserving strong robustness to harmful persona variation while enabling faithful persona adoption when required. Concretely, PerMix-RLVR improves persona stability score (PSS) over RLVR by +21.2% on MATH500, while also enhancing persona fidelity by +11.4% on PersonaGym.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.