2604.12177v1 Apr 14, 2026 cs.AI

LLM 기반 에이전트에서의 정책적 제약 무시 현상

Policy-Invisible Violations in LLM-Based Agents

J. Wu
J. Wu
Citations: 6
h-index: 2
Ming Gong
Ming Gong
Citations: 28
h-index: 2

LLM 기반 에이전트는 구문적으로 유효하고, 사용자로부터 승인을 받았으며, 의미적으로 적절한 행동을 수행할 수 있지만, 의사 결정 시 필요한 정보가 숨겨져 있어 조직의 정책을 위반할 수 있습니다. 우리는 이러한 현상을 '정책적 제약 무시 현상'이라고 부릅니다. 이는 준수 여부가 에이전트가 인지하는 맥락에 존재하지 않는 개체 속성, 상황 정보 또는 세션 기록에 따라 달라지는 경우를 의미합니다. 본 연구에서는 8가지 위반 유형을 포함하는 벤치마크인 PhantomPolicy를 제시합니다. PhantomPolicy는 위반 사례와 안전 사례를 균형 있게 포함하며, 모든 도구 응답에는 정책 메타데이터가 포함되지 않은 깔끔한 비즈니스 데이터만 포함되어 있습니다. 5개의 최첨단 모델이 생성한 600개의 모델 추적 데이터를 사람이 직접 검토하고, 사람이 검토한 추적 레이블을 사용하여 평가했습니다. 수동 검토 결과, 원래 사례 수준 주석과 비교하여 32개의 레이블(5.3%)이 변경되었으며, 이는 추적 수준의 수동 검토의 필요성을 확인시켜 줍니다. 긍정적인 조건 하에서 '세계 상태 기반 강제'가 달성할 수 있는 잠재력을 보여주기 위해, 우리는 반사실적 그래프 시뮬레이션을 기반으로 하는 강제 프레임워크인 Sentinel을 소개합니다. Sentinel은 모든 에이전트 행동을 조직 지식 그래프에 대한 제안된 변경 사항으로 취급하고, 행동 후의 세계 상태를 실현하기 위해 추론 실행을 수행하며, 그래프 구조적 불변성을 검증하여 허용/차단/명확화 결정을 내립니다. 사람이 검토한 추적 레이블을 기준으로 Sentinel은 콘텐츠 기반 DLP 기준(68.8% vs. 93.0% 정확도)보다 훨씬 우수한 성능을 보이며, 높은 정밀도를 유지하지만, 특정 위반 유형에 대한 개선 여지는 여전히 존재합니다. 이러한 결과는 정책 관련 세계 상태가 강제 계층에 제공될 때 달성할 수 있는 잠재력을 보여줍니다.

Original Abstract

LLM-based agents can execute actions that are syntactically valid, user-sanctioned, and semantically appropriate, yet still violate organizational policy because the facts needed for correct policy judgment are hidden at decision time. We call this failure mode policy-invisible violations: cases in which compliance depends on entity attributes, contextual state, or session history absent from the agent's visible context. We present PhantomPolicy, a benchmark spanning eight violation categories with balanced violation and safe-control cases, in which all tool responses contain clean business data without policy metadata. We manually review all 600 model traces produced by five frontier models and evaluate them using human-reviewed trace labels. Manual review changes 32 labels (5.3%) relative to the original case-level annotations, confirming the need for trace-level human review. To demonstrate what world-state-grounded enforcement can achieve under favorable conditions, we introduce Sentinel, an enforcement framework based on counterfactual graph simulation. Sentinel treats every agent action as a proposed mutation to an organizational knowledge graph, performs speculative execution to materialize the post-action world state, and verifies graph-structural invariants to decide Allow/Block/Clarify. Against human-reviewed trace labels, Sentinel substantially outperforms a content-only DLP baseline (68.8% vs. 93.0% accuracy) while maintaining high precision, though it still leaves room for improvement on certain violation categories. These results demonstrate what becomes achievable once policy-relevant world state is made available to the enforcement layer.

1 Citations
0 Influential
1 Altmetric
6.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!