2608.01676v1 Aug 03, 2026 cs.CL

대규모 언어 모델의 긴 문맥 처리에서 희소 어텐션 선택성의 인과 관계 분석: 반사실적 평가를 통한 이해

Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation

Haizhao Yang
Haizhao Yang
Citations: 1
h-index: 1
Xingyu Ren
Xingyu Ren
Citations: 41
h-index: 3
Youran Sun
Youran Sun
Citations: 5
h-index: 1
Chugang Yi
Chugang Yi
Citations: 6
h-index: 2

희소 어텐션은 긴 문맥을 처리하는 시스템에서 널리 사용되지만, 특정 블록을 버리는 것이 모델 출력에 미치는 영향에 대한 체계적인 분석은 부족했습니다. 본 연구에서는 먼저 희소 어텐션 선택성이 실제로 존재하며 인과 관계를 가진다는 것을 입증합니다. Block Sparse Flash Attention (BSFA)의 경로 재구현을 통해 4개의 아키텍처에서 16개 셀 중 13개에서 출력 결정이 변경되었으며, 정답/오답 레이블 전환은 전혀 발생하지 않았습니다. 그 후, 정밀하게 설계된 반사실적 감사 프레임워크를 도입합니다. 이 프레임워크는 Gold (정답 레이블 포함), Poison (의도적으로 잘못된 레이블 포함), 그리고 Benign (일반적인 내용) 블록을 사용하여 6가지 레이아웃 위치 대칭 조건 하에서, 희소화에 특화된 효과를 분리하여 분석합니다. 두 가지 상반되는 패턴이 관찰됩니다. 첫째, 신호 집중 현상으로 인해 선택자는 Gold 및 Poison 블록을 Benign 블록보다 훨씬 더 많이 유지하며 (모든 모델-태스크 쌍에서 G ≈ P ≫ B). 둘째, 통합 손실은 희소화 과정에서 블록 간의 어텐션 연결을 끊으며, 이는 probe block을 격리하면 해당 블록의 영향력이 4.48 logits에서 0으로 감소하는 ablation 실험을 통해 확인되었습니다. 압축 비율이 이 균형에 중요한 영향을 미칩니다. 4개의 모델-태스크 쌍에 대해 약한 ($c=0.25$) 압축부터 강력한 ($c=0.75$) 압축까지 다양한 수준에서 실험한 결과, 4개 셀 중 3개가 더 높은 압축률에서 희소 증폭 효과가 강해지는 경향을 보였으며, 2개의 셀에서는 극성 반전이 나타났습니다. BSFA 경로 재구현, 제어된 블록-top-$k$ 선택, 그리고 KV-캐시 제거라는 세 가지 독립적인 방법으로 실험한 결과, 희소화는 집계 정확도만으로는 감지할 수 없는 방식으로 콘텐츠의 영향력을 변화시킨다는 것을 알 수 있습니다. 본 연구에서는 모든 모델에서 블록 정보를 노출하는 측정 프레임워크를 제공합니다.

Original Abstract

Sparse attention is widely deployed in long-context serving stacks, yet no framework audits how discarding blocks changes the influence of specific content on model output. We first establish that the phenomenon is real and causal: Block Sparse Flash Attention (BSFA) route replay across four architectures changes output decisions in 13 of 16 cells, with zero identity-replay label flips. We then introduce a dense-calibrated counterfactual audit using matched probe cards---Gold (carrying the correct answer label), Poison (carrying a target wrong label), and Benign (filler only)---under six-layout position symmetry, isolating the sparsification-specific effect. Two patterns compete. Signal concentration: the selector preserves Gold and Poison blocks far above filler-matched Benign blocks (G$\approx$P$\gg$B across all model--task pairs). Integration loss: discarding blocks severs cross-block attention---confirmed by an ablation where isolating the probe block collapses its influence from 4.48 logits to zero. Compression ratio governs the balance: a full sweep from mild ($c=0.25$) to aggressive ($c=0.75$) compression across four model--task pairs reveals that three of four cells move toward stronger sparse amplification at higher compression, with two exhibiting sign reversals. Three independent arms---BSFA route replay, controlled block-top-$k$, and KV-cache eviction---converge: sparsification changes content influence in ways aggregate accuracy cannot detect. We provide an open measurement framework deployable on any model exposing block identities.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!