2607.20827v1 Jul 23, 2026 cs.AI

LLM 에이전트 액션 선택 시 출처 정보의 민감도 감사

Auditing Provenance Sensitivity in LLM Agent Action Selection

Jun-Hui Liao
Jun-Hui Liao
Citations: 4
h-index: 1

LLM 에이전트는 사용자 요청, 도구 출력, 검색된 기록, 메모리 및 신뢰할 수 없는 텍스트가 혼합된 컨텍스트에서 도구와 인수를 선택합니다. 증거는 의사 결정에 허용되지 않더라도 관련될 수 있으므로, 올바른 액션은 반드시 허용된 증거만을 기반으로 해야 합니다. 본 연구에서는 각 도구 및 인수 대상에 대해 컨텍스트 요소를 개별적으로 라벨링하는 대상별 권한 감사 방식을 소개합니다. 주요 실험에서 작업, 명제, 위치 및 정책을 고정한 상태로 명제의 출처 권한만 변경합니다. 또한 유효한 증거가 약화될 때의 동작을 테스트하고, 컨텍스트 부분 집합 상호 작용을 2차 로컬화 진단으로 사용합니다. 450개의 제어된 다음 액션 작업과 여러 개의 오픈 웨이트 LLM 패밀리에 대해, 신뢰할 수 있는 변형과 신뢰할 수 없는 변형은 경쟁적인 경우의 5.4%에서 서로 다른 액션을 수행하는 반면, 지원적인 경우에는 1.7%에서 차이를 보입니다. 제어된 성능 저하 조건 하에서는, 권한이 없는 경쟁이 전체적으로 올바른 결과, 혼합 오류 및 깨끗한 결과를 보이는 패턴으로 나타나며, 이는 2.4%의 비교에서 관찰됩니다 (95% 신뢰 구간: 2.1% ~ 3.0%). 이러한 수치는 배포 환경에서의 일반적인 발생률이 아닌, 통제된 스트레스 테스트의 결과입니다. 모델은 텍스트 기반 출처 권한 단서에 반응하지만, 이는 신뢰할 수 없는 증거가 액션에 영향을 미치지 못하게 하는 것을 방지하지 않습니다.

Original Abstract

LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action need not be grounded only in permitted evidence. We introduce a target-specific authorization audit that labels context factors separately for each tool and argument target. Its primary test holds the task, proposition, position, and policy fixed while changing only the proposition's source authority. We then test behavior when valid evidence is weakened and use context-subset interactions as a secondary localization diagnostic. Across 450 controlled next-action tasks and multiple open-weight LLM families, trusted and untrusted variants produce different actions in 5.4 percent of competing cases versus 1.7 percent of supporting cases. Under controlled degradation, unauthorized competition is retained in a full-correct, mixed-error, clean-correct pattern in 2.4 percent of comparisons, with a 95 percent confidence interval from 2.1 to 3.0 percent. These are controlled stress-set rates, not deployment prevalence. The models respond to textual source-authority cues, but this does not prevent untrusted evidence from influencing their actions.

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!