그래프 어텐션은 언제 희소해야 하는가? 각 엣지별 Tsallis 지수 학습
When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index
그래프 어텐션은 소프트맥스 함수를 사용하여 이웃 노드의 중요도를 정규화하는데, 이는 Shannon 통계 하에서의 최대 엔트로피 방식입니다. 하지만 동종성 그래프와 이질성 그래프는 서로 다른 형태의 어텐션을 필요로 하며, 하나의 고정된 정규화 방식으로는 두 경우 모두에 적합하지 않습니다. 본 논문에서는 가중치와 함께 Tsallis 엔트로피 지수 $q$를 공동으로 학습하는 그래프 어텐션 레이어인 **LTGA** (**L**earnable **T**sallis **G**raph **A**ttention)을 제안합니다. LTGA는 전역 스칼라 값부터 각 엣지별 지수까지, 네 가지 수준에서 heavy-tailed ($q<1$), softmax ($q=1$) 및 compact-support ($q>1$) 어텐션을 연속적으로 보간하며, bounded reparameterization을 통해 모든 모델이 GAT 기본 설정을 따르도록 합니다. 8개의 벤치마크와 10개의 실험 설정에서 LTGA-Edge는 가장 좋은 평균 순위($2.75$)를 기록했지만, 전반적인 검정 결과에서는 통계적 유의미성이 관찰되지 않았습니다 ($p=0.199$). 또한, $q$ 값을 학습하는 것이 최적값을 탐색하는 것보다 성능이 좋지 않다는 것을 확인했습니다. 검증 데이터를 사용하여 튜닝된 고정 그리드 방식은 $61.4%$, 튜닝된 α-entmax는 $62.2%$, 그리고 동일한 용량을 가진 $q=1$인 제어 그룹은 $62.0%$의 성능을 보였으며, LTGA-Edge는 $61.7%$를 기록했습니다. 학습된 지수가 제공하는 장점은 그리드 탐색 대신 한 번의 실행으로 결과를 얻을 수 있다는 점과 해석 가능한 메커니즘을 제공한다는 것입니다. $q$ 값이 1일 때, LTGA는 어텐션 계수의 42%를 정확히 0으로 만들어서 불필요한 연결을 제거하며, 이러한 연결을 복구하는 데에는 7.1점이 손실됩니다. 반면, 동일 비율로 무작위적으로 가지치기를 수행하면 13.0점이 더 손실됩니다. 프로젝트 페이지: https://kleyt0n.github.io/ltga
Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention shapes, and one fixed normalization cannot serve both. We propose \textbf{LTGA} (\textbf{L}earnable \textbf{T}sallis \textbf{G}raph \textbf{A}ttention), a graph attention layer whose Tsallis entropic index $q$ is learned jointly with the weights, interpolating continuously between heavy-tailed ($q\!<\!1$), softmax ($q\!=\!1$) and compact-support ($q\!>\!1$) attention at four granularities from a global scalar to a per-edge index, under a bounded reparameterization that starts every model at the GAT baseline. Across eight benchmarks at ten seeds, LTGA-Edge takes the best average rank ($2.75$), but the omnibus test does not reject ($p\!=\!0.199$) and learning $q$ does not beat searching it: a validation-tuned frozen grid reaches $61.4\%$, tuned $α$-entmax $62.2\%$ and a capacity-matched $q\!\equiv\!1$ control $62.0\%$, against $61.7\%$ for LTGA-Edge. What the learned index buys is one run instead of a grid, and an interpretable mechanism: where $q$ leaves $1$, it prunes $42\%$ of attention coefficients to exactly zero, and those edges are selectively the wrong ones, restoring them costs $7.1$ points, while random pruning at the same rate costs $13.0$ more. Project page: https://kleyt0n.github.io/ltga
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.