2607.01557v1 Jul 02, 2026 cs.CL

DiPS: 고위험 설득 에이전트를 위한 대화 정책 선택

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents

Abrar Anwar
Abrar Anwar
Citations: 385
h-index: 9
Tianyi Zhang
Tianyi Zhang
Citations: 79
h-index: 4
Mousumi Das
Mousumi Das
Citations: 0
h-index: 0
Jesse Thomason
Jesse Thomason
Citations: 83
h-index: 4
David Traum
David Traum
Citations: 5
h-index: 2

대규모 언어 모델(LLM)은 종종 고위험 상황에서의 설득에 어려움을 겪습니다. 사람들의 개별적인 성격과 우려 사항은 획일적인 접근 방식보다는 맞춤형 전략을 요구합니다. 이러한 과제를 해결하기 위해, 우리는 화재 및 구조 시나리오를 연구하며, 연소자를 대피시키도록 설득하는 고위험 설득 영역에서 '대화 정책 선택(DiPS)'이라는 Q-러닝 프레임워크를 제안합니다. 구체적으로, 우리는 비평 네트워크를 훈련시켜 대화의 맥락 변화에 맞춰 적절한 설득 전략을 동적으로 선택하도록 합니다. 이 비평 네트워크는 연소자의 최근 발언을 기반으로 각 단계에서 최적의 설득 정책을 선택하는 역할을 합니다. 우리는 DiPS를 시뮬레이션 환경과 실제 사람과의 상호 작용 모두에서 다양한 기준 모델과 비교하여 평가했습니다. 그 결과, DiPS는 제로샷 LLM 및 일반적인 RAG(Retrieval-Augmented Generation) 기반 접근 방식보다 더 높은 대피 성공률을 달성하는 것으로 나타났습니다.

Original Abstract

Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns require tailored strategies rather than a one-size-fits-all approach. To address this challenge, we focus on a fire-rescue scenario in which an operator must persuade a resident to evacuate as a high-stakes persuasion domain and propose Dialogue Policy Selection (DiPS), a Q-learning framework to dynamically select persuasion strategies adapted to the evolving conversational context. Specifically, we train a critic, trained to maximize the chance of evacuation success, to select a persuasion policy at each turn based on the resident's recent utterances.We then evaluate DiPS against multiple baselines in both simulated and real human interactions. We find that DiPS achieves higher evacuation success than a zero-shot LLM and generic RAG-augmented approach.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!