복잡한 행동 모델링: 비전-언어 모델에서의 다중 인격 구성 및 동적 전환
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models
다중 모드 대규모 언어 모델(MLLM)이 사회적 상호 작용에 광범위하게 사용됨에 따라, 복잡한 성격 조건 하에서 이들의 행동을 이해하고 제어하는 것이 필수적입니다. 본 논문에서는 명시적인 성격 조건을 부여하고, 단일 성격 유도, 다중 성격 유도 및 성격 전환을 포괄하는 체계적인 평가 프레임워크를 제시합니다. 실험 결과, 성격 유도는 이미지 캡셔닝 성능을 향상시키지만, 시각적 질문 응답(VQA)과 같이 정밀한 추론이 필요한 작업의 성능을 저하시킬 수 있음을 보여줍니다. 다중 특성 구성 및 동적 전환 과정에서 균형 문제와 잔류 효과가 관찰되었으며, 이는 모델의 행동이 이전 및 현재 성격 제약 모두에 의해 공동으로 조절됨을 시사합니다. 기존의 프롬프트 기반 성격 유도 방법은 다중 모드 환경으로의 적용 가능성이 제한적입니다. 본 연구는 MLLM에서 성격 모델링의 역동적이고 복잡한 특성을 밝히고, 성격 유도 및 평가를 위한 견고하고 맞춤화된 방법에 대한 필요성을 강조합니다. 논문이 게재되면 코드를 공개할 예정입니다.
With the widespread deployment of Multimodal Large Language Models (MLLMs) in social interaction, understanding and controlling their behavior under complex personality conditions is essential. This paper introduces explicit personality conditioning and establishes a systematic evaluation framework encompassing single-personality induction, multi-personality induction, and personality switching. Experiments show that personality induction improves image captioning performance but can impair performance on tasks requiring precise reasoning, such as visual question answering (VQA). Balancing and residual effects are observed during multi-trait composition and dynamic switching, indicating that model behavior is co-modulated by both previous and current personality constraints. Existing prompt-based personality induction methods show limited transferability to multimodal settings. Our work reveals the dynamic and complex nature of personality modeling in MLLMs and underscores the need for robust, tailored methods for personality induction and evaluation. The code will be released when the paper is accepted.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.