PHASE-Tree: 장기적인 역할극 대화에서 캐릭터 상태의 변화 모델링
PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue
장기적인 역할극에서는 캐릭터가 스토리 전개에 따라 일관성을 유지해야 합니다. 그러나 기존 연구는 다음과 같은 두 가지 측면에서 부족합니다. 첫째, 대부분의 표현 방식은 정적인 프로필로 구성되어 있으며, 변경되지 않은 특성을 불안정하게 만들지 않고도 특정 부분만 업데이트하기 어렵습니다. 둘째, 벤치마크는 주로 페르소나 유지 및 기억력 회복을 테스트하는 데 중점을 두며, 모델이 캐릭터의 현재 발전된 상태에 맞춰 응답하는지를 평가하지 않습니다. 본 논문에서는 이러한 문제점에 대한 해결책을 제시합니다. PHASE-Tree는 불변의 핵심 아이덴티티를 가지는 다중 시간 척도 기반의 캐릭터 상태 트리이며, 페르소나, 세션 및 순간 레이어를 통해 각 변경 가능한 필드를 개별 에피소드 내외부의 업데이트를 위한 대상 요소로 활용할 수 있습니다. PHASE-Tree는 명시적인 텍스트 제공 또는 암시적 파라미터 적응을 통해 응답을 생성합니다. 캐릭터의 발전된 상태에 따른 응답 생성을 평가하기 위해, 본 논문에서는 네 개의 장기 대화 코퍼스를 사용하여 에피소드 간 발전을 측정하고, 네 개의 짧은 대화 코퍼스를 사용하여 장면 내 상태 추적을 확인하는 LongEvoRoleBench를 제안합니다. 장기 대화 데이터셋에서 PHASE-Tree는 자체 변형 모델 및 외부 텍스트 기반 모델과 비교하여 12개의 데이터셋-메트릭 조합 중 11개에서 가장 높은 순위를 기록했으며, 캐릭터 수준, 의미론적, 임베딩 점수를 각각 19.7%, 12.4%, 15.1% 향상시켰습니다. 블라인드 테스트에서는 인간 평가가 GPT-4.1 판별기의 결과와 상관 관계를 보였으며 (Pearson r = 0.65), 설명적인 n=10의 프롬프트 서브셋에서 평균 점수 차이는 +0.20이었습니다. 이러한 장기 대화에서의 성능 우위는 다양한 LLM 판별기와 생성 모델 아키텍처에서도 유지되었습니다.
Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts: representations are typically static profiles that cannot be updated locally without destabilizing unchanged traits, and benchmarks mainly test persona preservation and memory recall rather than whether a model speaks from a character's currently evolved state. We address both. PHASE-Tree is a multi-timescale character-state tree with an immutable identity root and mutable persona, session, and moment layers, making each mutable field an addressable target for localized within- and cross-episode updates. It conditions generation through explicit textual provision or implicit parametric adaptation. To measure evolved-state generation, we introduce LongEvoRoleBench, which pairs four long-dialogue corpora for cross-episode evolution with four short-dialogue corpora as within-scene state-tracking checks, under a unified next-utterance protocol. On the long-dialogue core, textual PHASE-Tree ranks first in 11 of 12 dataset-metric cells against internal variants and all 12 cells against external textual baselines, improving character-level, semantic, and embedding scores by 19.7%, 12.4%, and 15.1% respectively. In a blinded 200-response study, human ratings correlate with the GPT-4.1 judge (Pearson r= 0.65); on descriptive n= 10 PT and NR prompt subsets, the Overall difference is +0.20. The long-dialogue Sem advantage persists across LLM judges and generation backbones.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.