계층적 추론과 발화 수준의 목표 보상을 통한 LLM의 사회 지능 향상
Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding
대규모 언어 모델(LLM)은 구조화된 작업에서는 뛰어난 성능을 보이지만, 장기적인 목표 조정 및 빠른 적응이 필요한 역동적인 사회적 상호작용에는 어려움을 겪습니다. 기존 방법들은 종종 모든 발화에 동일한 목표 기반 보상을 적용하여, 각 대화 단계에서의 구체적인 목표를 간과하고 잠재적인 전략의 타당성을 고려하지 못합니다. 본 연구는 계획된 행동 이론(Theory of Planned Behavior)에서 영감을 받아, 사회적 대화를 고차원의 전략 수립과 저차원의 언어적 실행이라는 두 가지 계층 구조로 분해하는 Think-Strategy-Response (TSR) 프레임워크를 제안합니다. TSR을 최적화하기 위해, 목표 달성 점수의 변동성에 따라 보상을 동적으로 조정하여 목표 완료와 전략 준수 사이의 균형을 맞추는 새로운 알고리즘인 Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR)를 도입했습니다. SOTOPIA 벤치마크에서의 실험 결과, 제안하는 방법은 Qwen2.5-7B 에이전트를 미세 조정하여 GPT-4o의 기준 성능을 7.32% 이상 능가하며, 다중 에이전트 사회적 협상 작업에서 최첨단 성능을 달성함을 보여줍니다.
Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and failing to account for the rationale of potential strategies. Inspired by the Theory of Planned Behavior, we propose the Think-Strategy-Response (TSR) framework, which decomposes social dialogue into two hierarchical stages: high-level strategic planning and low-level linguistic execution. To optimize TSR, we introduce Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR), a novel algorithm that dynamically routes rewards - balancing goal completion and strategy adherence - based on the variance of goal achievement scores. Experiments on the SOTOPIA benchmark show that our approach fine-tunes a Qwen2.5-7B agent to surpass the GPT-4o baseline by 7.32% in goal completion success, demonstrating state-of-the-art performance in multi-agent social negotiation tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.