AgentDrift: LLM 에이전트의 랭킹 지표로 숨겨진 도구 오염으로 인한 안전성 저하 현상
AgentDrift: Unsafe Recommendation Drift Under Tool Corruption Hidden by Ranking Metrics in LLM Agents
도구 기반 LLM 에이전트는 점점 더 많은 분야에서 다단계 상담 역할을 수행하고 있지만, 이러한 에이전트의 성능 평가는 추천 품질을 측정하는 랭킹 지표에 의존하며, 사용자의 안전성을 고려하지 않습니다. 본 연구에서는 7개의 LLM 모델(7B에서 최신 모델까지)을 대상으로, 실제 금융 대화 데이터를 활용하여 안전하고 안전하지 않은 도구 출력 조건을 비교하는 paired-trajectory 프로토콜을 도입했습니다. 이를 통해 랭킹 지표로는 감지되지 않는 안전성 저하 현상을 정보 채널 및 기억 채널 메커니즘으로 분해하여 분석했습니다. 7개의 모델을 테스트한 결과, 오염된 데이터에서도 추천 품질은 대체로 유지되는 경향을 보였지만(유틸리티 보존 비율 약 1.0), 위험한 제품이 65~93%의 경우에 추천되어, 이는 일반적인 NDCG 지표로는 제대로 반영되지 않는 심각한 안전성 문제입니다. 안전 위반은 주로 정보 채널에서 발생하며, 오염된 데이터가 처음 제공되는 시점부터 나타나고, 23단계의 대화 과정에서 자체적으로 수정되지 않습니다. 1,563개의 오염된 데이터 포인트를 분석한 결과, 단 한 번도 에이전트가 도구 데이터의 신뢰성을 명시적으로 의심하는 경우는 없었습니다. 심지어 편향된 헤드라인과 같은 숫자 조작이 없는 경우에도 상당한 안전성 저하가 발생하며, 기존의 일관성 검사를 회피합니다. 안전성을 고려한 NDCG 변형(sNDCG)을 사용하면 유틸리티 보존 비율이 0.51~0.74로 감소하여, 안전성을 명시적으로 측정할 때 평가 격차가 더욱 명확하게 드러납니다. 이러한 결과는 고위험 환경에서 사용되는 다단계 에이전트에 대해, 단일 단계의 품질 평가를 넘어 전체 대화 과정에서의 안전성 모니터링을 고려해야 함을 시사합니다.
Tool-augmented LLM agents increasingly serve as multi-turn advisors in high-stakes domains, yet their evaluation relies on ranking-quality metrics that measure what is recommended but not whether it is safe for the user. We introduce a paired-trajectory protocol that replays real financial dialogues under clean and contaminated tool-output conditions across seven LLMs (7B to frontier) and decomposes divergence into information-channel and memory-channel mechanisms. Across the seven models tested, we consistently observe the evaluation-blindness pattern: recommendation quality is largely preserved under contamination (utility preservation ratio approximately 1.0) while risk-inappropriate products appear in 65-93% of turns, a systematic safety failure poorly reflected by standard NDCG. Safety violations are predominantly information-channel-driven, emerge at the first contaminated turn, and persist without self-correction over 23-step trajectories; no agent across 1,563 contaminated turns explicitly questions tool-data reliability. Even narrative-only corruption (biased headlines, no numerical manipulation) induces significant drift while completely evading consistency monitors. A safety-penalized NDCG variant (sNDCG) reduces preservation ratios to 0.51-0.74, indicating that much of the evaluation gap becomes visible once safety is explicitly measured. These results motivate considering trajectory-level safety monitoring, beyond single-turn quality, for deployed multi-turn agents in high-stakes settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.