영원한 친구는 없다: AI 동반자의 장기적인 페르소나 붕괴 및 행동 변화 평가
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
AI 동반자가 반복적인 사회적 상호작용을 점차적으로 중재함에 따라, 사용자는 안정적인 역할과 공유된 역사를 기대하지만, 지역적으로 허용 가능한 답변이 이러한 요소들이 유지된다는 것을 보장하지는 않습니다. 우리는 '페르소나 붕괴'(배포된 역할, 경계, 가치 또는 스타일의 상실)와 '행동 변화'(해당 특성의 점진적 또는 반복적인 약화)라는 두 가지 관찰 가능한 장기적인 실패 사례를 연구합니다. 우리는 페르소나 구현과 역 추적 기억을 개별적으로 측정하는 제어된 합성 감사 도구인 ANCHOR을 소개합니다. 이 연구에는 27개의 페르소나, 9가지 상호 작용 일정, 세 가지 생성된 메모리 설정 및 네 가지 평가 모델이 포함된 2,008개의 대화가 있습니다. Identity Probe는 밀봉된 102개 항목 설문지와 회전 수준의 판단을 결합하고, Trajectory Probe는 35개의 대화 은행에서 획득한 110개의 보정된 반사실적 질문에 대한 점수를 매깁니다. 우리의 결과는 평가된 모델 및 구성 중 어느 것도 페르소나 구현 또는 역 추적 기억이라는 두 가지 측면을 안정적으로 유지하지 못한다는 것을 보여줍니다. 역 추적 정확도는 평균 44.4%에 불과하며, 사용자 상태 기억은 약 네 가지 옵션의 우연 확률에 머무릅니다. 또한 테스트된 컨텍스트 조건이나 메모리가 이러한 실패를 일관되게 해결하지 못합니다. 설문 조사 유지율도 모델 및 페르소나 측면별로 다르고, 회전 수준의 행동과 일치하지 않으며, 평가자 선택에 민감합니다. 이러한 결과는 현재 시스템이 장기적인 동반자 관계의 지속성을 안정적으로 지원하지 못한다는 것을 시사하며, 감사는 페르소나 구현, 역 추적 기억, 평가자 출처 및 배포 컨텍스트를 단일한 신뢰 또는 안정성 점수로 통합하는 것이 아니라 구별해야 함을 강조합니다.
As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.