2608.12935v1 Aug 13, 2026 cs.AI

섭동 반응에서의 증거 분해, 모순 및 취약성 분석

Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

Lei You
Lei You
Citations: 50
h-index: 4

섭동 방법은 입력값을 변경하여 모델의 예측 변화를 측정함으로써 모델의 의사 결정을 설명합니다. 그러나 반응의 크기는 모델이 얼마나 반응하는지를 나타낼 뿐이며, 그 반응의 의미에 대한 정보는 제공하지 않습니다. 동일한 크기의 반응이라도 최종적인 사실과 반사실 간의 차이를 뒷받침하거나, 반대로 이에 반대할 수 있으며, 섭동 경로를 따라 강하게 나타났다가 최종 지점에서 사라질 수도 있습니다. 따라서 본 연구에서는 쌍을 이루는 입력값이 점진적으로 드러나면서 대비(contrast)가 어떻게 변화하는지를 추적하고, 최종 대비를 사용하여 이러한 변화 과정을 해석합니다. 우리는 증거(E), 모순(C), 그리고 취약성(F)을 분해하는 방법인 DECAF (Decomposition of Evidence, Contradiction, And Fragility)을 제시합니다. DECAF은 정렬된 반응, 반대되는 반응, 그리고 최종 지점에서 사라지는 반응들을 각각 증거 E, 모순 C, 그리고 취약성 F로 분류합니다. 이러한 분해는 일반적인 크기(Abs)를 정확하게 보존하며, Abs = E + C + F의 관계를 만족합니다. 또한, DECAF은 독립적으로 측정된 행동과 세 가지 구성 요소가 서로 연관성을 가짐을 보여줍니다. 72개의 ImageNet-9 모델에 대한 분석에서는 거의 동일한 반응 크기를 갖지만, 독립적으로 측정된 행동이 다른 사례들을 비교했습니다. 가장 큰 DECAF 구성 요소는 관찰된 행동과 96.4%의 경우에 일치하는 반면, 크기만으로는 35.0%의 일치율을 보였습니다. 입력값 공개 경로를 변경하는 것만으로 전체 반응이 약 80% 증가하지만, 증거는 거의 변하지 않으며 취약성은 4배 이상 증가합니다. FunnyBirds 및 ImageNet-1k 데이터셋에서 짧은 DECAF 추세(trajectory)는 기존의 일반적인 설명 기법들보다 우수한 성능을 보였습니다. 1B 규모의 DINOv2 모델에서, 짧은 DECAF 추세는 강력한 경사 기반 방법과 비교하여 계산 시간은 4.75배 단축되고 메모리 사용량은 2.36배 감소했습니다.

Original Abstract

Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We introduce DECAF (Decomposition of Evidence, Contradiction, And Fragility), which routes aligned, opposed, and endpoint-null responses into evidence E, contradiction C, and fragility F. The decomposition preserves ordinary magnitude exactly, Abs = E + C + F, and is unique under endpoint-relative axioms. Across controlled vision and tabular settings, the three components track independently measured behavior. In a 72-model ImageNet-9 audit, we compare cases with nearly identical response magnitude but different independently measured behaviors. The largest DECAF component agrees with an observed behavior in 96.4% of cases, compared with 35.0% for magnitude alone. Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4x. On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform the tested general-purpose attribution baselines. On a 1B-scale DINOv2 model, a short trajectory matches a strong gradient-based baseline with 4.75x lower wall time and 2.36x lower peak memory.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!