2605.27157v1 May 26, 2026 cs.AI

탐지는 해결이 아니다: 검색 증강 LLM에서의 모니터링 제어 격차

Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Changting Lin
Changting Lin
Citations: 216
h-index: 8
Wenpeng Xing
Wenpeng Xing
Citations: 205
h-index: 10
Chenchen Ye
Chenchen Ye
Citations: 66
h-index: 3
Zhengtao Yu
Zhengtao Yu
Citations: 20
h-index: 2
Xuyang Teng
Xuyang Teng
Citations: 6
h-index: 2
Meng Han
Meng Han
Citations: 78
h-index: 3

검색 증강 LLM은 증거의 품질이 행동의 안전성을 결정하는 작업에 사용되지만, 기존 평가 프로토콜은 단일 턴(turn)의 강건성이 여러 턴에 걸쳐 증거가 축적될 때의 강건성을 예측한다고 가정합니다. 본 연구에서는 이러한 가정이 근본적으로 잘못되었음을 보여줍니다. 모델들은 모니터링과 제어 사이의 격차를 보입니다. 즉, 모델은 상반되는 증거를 쉽게 인지하지만, 이러한 인식이 최종 권장 사항을 제한하지 못한다는 것입니다. 지식적 충돌(epistemic conflict)을 감지하는 것이 반드시 안전하게 해결하는 것을 의미하지 않습니다. 본 연구에서는 1.5B에서 32B 파라미터 사이의 네 가지 모델 패밀리에 대한 다중 턴 문서 축적 프로토콜과 5만 건 이상의 턴 단위 평가를 통해, 단일 턴 진단이 RAG 안전성을 체계적으로 과대평가하고, 상반되는 내용에 대한 인지가 안전한 해결과 상관관계가 없으며, 이러한 경향은 표적 인간 검증을 통해 뒷받침된다는 것을 입증했습니다. 또한 보편적인 프롬프트 수정 방법은 존재하지 않습니다. 은닉 상태 탐색, 어텐션 분석 및 응답 전략 분류를 포함하는 다양한 메커니즘 증거는 행동 선택이 문제의 핵심 영역임을 시사합니다. 위험과 관련된 정보가 내부적으로 표현되고 안전하지 않은 생성 과정에서 더 많은 관심을 받지만, 출력 동작을 제한하는 데 실패합니다. 모델이 인식하는 것과 실제로 수행하는 것 사이의 격차를 측정하고 해소해야 검색 증강 시스템이 고위험 환경에서 신뢰할 수 있게 됩니다.

Original Abstract

Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robustness predicts robustness when evidence accumulates across turns. We show this assumption is fundamentally incorrect. Models exhibit a monitoring-control gap: they readily acknowledge contradictory evidence, yet this awareness fails to constrain their final recommendations - detecting epistemic conflict does not imply resolving it safely. Through a multi-turn document accumulation protocol across four model families (1.5B-32B parameters) and over 50,000 turn-level evaluations, we demonstrate that single-turn diagnostics systematically overestimate RAG safety, that contradiction acknowledgement is uncorrelated with safe resolution, a pattern corroborated by targeted human validation, and that no universal prompt fix exists. Converging mechanism evidence - hidden-state probing, attention analysis, and response-strategy taxonomy - points to action selection as the most plausible locus of the deficit: danger-relevant information is internally represented and receives enhanced attention during unsafe generation, yet fails to constrain output behavior. The gap between what models recognize and what they do must be measured and closed before retrieval-augmented systems can be trusted in high-stakes settings.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!