DiPS: 고위험 설득 에이전트를 위한 대화 정책 선택
DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents
대규모 언어 모델(LLM)은 종종 고위험 상황에서의 설득에 어려움을 겪습니다. 사람들의 개별적인 성격과 우려 사항은 획일적인 접근 방식보다는 맞춤형 전략을 요구합니다. 이러한 과제를 해결하기 위해, 우리는 화재 및 구조 시나리오를 연구하며, 연소자를 대피시키도록 설득하는 고위험 설득 영역에서 '대화 정책 선택(DiPS)'이라는 Q-러닝 프레임워크를 제안합니다. 구체적으로, 우리는 비평 네트워크를 훈련시켜 대화의 맥락 변화에 맞춰 적절한 설득 전략을 동적으로 선택하도록 합니다. 이 비평 네트워크는 연소자의 최근 발언을 기반으로 각 단계에서 최적의 설득 정책을 선택하는 역할을 합니다. 우리는 DiPS를 시뮬레이션 환경과 실제 사람과의 상호 작용 모두에서 다양한 기준 모델과 비교하여 평가했습니다. 그 결과, DiPS는 제로샷 LLM 및 일반적인 RAG(Retrieval-Augmented Generation) 기반 접근 방식보다 더 높은 대피 성공률을 달성하는 것으로 나타났습니다.
Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns require tailored strategies rather than a one-size-fits-all approach. To address this challenge, we focus on a fire-rescue scenario in which an operator must persuade a resident to evacuate as a high-stakes persuasion domain and propose Dialogue Policy Selection (DiPS), a Q-learning framework to dynamically select persuasion strategies adapted to the evolving conversational context. Specifically, we train a critic, trained to maximize the chance of evacuation success, to select a persuasion policy at each turn based on the resident's recent utterances.We then evaluate DiPS against multiple baselines in both simulated and real human interactions. We find that DiPS achieves higher evacuation success than a zero-shot LLM and generic RAG-augmented approach.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.