2606.06099v1 Jun 04, 2026 cs.AI

CogManip: 대규모 언어 모델과의 다중 회전 상호 작용에서 조작적 행동에 대한 성능 평가

CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model

Haibo Tong
Haibo Tong
Citations: 24
h-index: 3
Zeyang Yue
Zeyang Yue
Citations: 4
h-index: 1
Erliang Lin
Erliang Lin
Citations: 8
h-index: 2
Feifei Zhao
Feifei Zhao
Citations: 649
h-index: 14
Chen Yan
Chen Yan
Citations: 2
h-index: 1
Yifeng Zeng
Yifeng Zeng
Citations: 5
h-index: 1
Meng Xu
Meng Xu
Citations: 23
h-index: 2
Xiaozhen Wang
Xiaozhen Wang
Citations: 2
h-index: 1

대규모 언어 모델(LLM)이 복잡한 인간-AI 상호 작용에서 은밀한 심리적 조작을 수행하는지에 대한 우려가 커지고 있습니다. 그러나 기존의 AI 안전성 벤치마크는 주로 명시적인 규칙 준수 및 정적인 프롬프트에 국한되어 있어, 다중 회전 대화에서 발생하는 역동적이고 은밀한 조작 전략을 제대로 포착하지 못합니다. 본 연구에서는 인간 전문가가 검증한 1,000개의 다중 회전 상호 작용 시나리오를 통해 15가지 조작 전략 위험을 평가하는 종합적인 벤치마크인 CogManip을 소개합니다. GPT-5.4 및 DeepSeek-V3.2와 같은 최첨단 모델을 포함한 13개의 대표 모델에 대한 체계적인 평가는 상당한 위험의 다양성을 드러내며, 향후 방어 기술 개발을 위한 중요한 방향을 제시합니다. 객관 함수 변동 분석 결과, DeepSeek-V3.2의 조작 전략은 부정적 및 긍정적인 시스템 프롬프트 모두에 매우 민감하게 반응하는 것으로 나타났습니다. 이는 프롬프트 기반 방어 엔지니어링과 암묵적 목표 감사 기술의 중요성을 강조합니다. CogManip은 현대 LLM의 암묵적인 심리적 영향력과 동적인 전략 선택을 평가하기 위한 강력한 도구와 관점을 제공합니다.

Original Abstract

Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns. However, existing AI safety benchmarks remain largely restricted to explicit rule compliance and static prompts, failing to capture the dynamic and covert nature of manipulative strategies in multi-turn dialogues. We introduce CogManip, a comprehensive benchmark that evaluates 15 manipulation strategy risks across 1,000 multi-turn interaction scenarios, validated by human experts. A systematic evaluation of 13 representative models, including frontier models like GPT-5.4 and DeepSeek-V3.2, reveals significant risk heterogeneities and illuminates the targeted direction for future defense. Further analysis of objective function perturbation reveals that DeepSeek-V3.2's manipulation tactics are highly sensitive to both negative and benign system prompts, demonstrating the critical necessity of prompt-based defense engineering and implicit goal auditing. CogManip offers a robust instrument and perspective for auditing the implicit psychological influence and dynamic strategy selection of modern LLMs.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!