2606.10650v1 Jun 09, 2026 cs.CL

동적 선형 어텐션

Dynamic Linear Attention

Bo Zheng
Bo Zheng
Citations: 44
h-index: 3
Xin Wang
Xin Wang
Citations: 7,232
h-index: 39
Xueshen Liu
Xueshen Liu
Citations: 58
h-index: 4
Hui Shen
Hui Shen
Citations: 402
h-index: 10
Zesen Zhao
Zesen Zhao
Citations: 41
h-index: 2
Minkyoung Cho
Minkyoung Cho
Citations: 29
h-index: 2
Zhongwei Wan
Zhongwei Wan
Citations: 1,260
h-index: 16
Z. Mao
Z. Mao
Citations: 30
h-index: 2
Shen Yan
Shen Yan
Citations: 452
h-index: 5
Mi Zhang
Mi Zhang
Citations: 692
h-index: 13

대규모 언어 모델(LLM)의 장문 맥락 처리 능력은 표준 어텐션의 이차 복잡성으로 인해 근본적으로 제한됩니다. 따라서, 더 낮은 복잡도를 갖는 선형 어텐션 기법이 널리 사용되고 있습니다. 최근 연구에서는 장문 맥락에서 표현 능력을 향상시키기 위해 메모리를 다중 상태로 구성하는 방법이 제시되었습니다. 그러나 기존의 다중 상태 선형 어텐션 방식은 고정된 상태 병합 정책에 의존하며, 이는 동적으로 변화하는 토큰의 중요도에 적응할 수 없어 중요한 토큰을 영구적으로 가리고 긴 시퀀스에서 심각한 오류를 누적시키는 문제를 야기합니다. 이러한 한계를 해결하기 위해, 우리는 다중 상태 선형 어텐션을 위한 동적 메모리 모델링 프레임워크인 DLA를 제안합니다. DLA는 (i) 정보 기반의 동적 상태 병합을 도입하여 토큰 수준의 정보 변화에 따라 상태 경계를 적응적으로 결정함으로써 의미 변환 영역에서는 고해상도 표현을 유지하고 안정적인 영역은 적극적으로 요약하며, (ii) 용량 제한 메모리 모델링을 통해 시계열 순서로 정렬된 고정 크기 상태 캐시를 유지하면서 인접한 낮은 정보 상태를 선택적으로 병합하여 최소한의 정보 손실로 메모리 증가를 제어합니다. 우리는 DLA를 두 가지 다른 선형 어텐션 모델에서 사전 훈련하고 세 가지 범주의 16개 데이터셋에 대해 평가했습니다. 실험 결과는 DLA가 최첨단 기술보다 우수함을 보여줍니다.

Original Abstract

The scalability of Large Language Models (LLMs) to long contexts is fundamentally constrained by the quadratic complexity of standard attention, motivating the adoption of linear attention mechanisms with sub-quadratic cost. To improve representation capacity under long contexts, recent approaches organize memory in a multi-state manner. However, existing multi-state linear attention methods rely on fixed state merging policies that cannot adapt to dynamically varying token importance, irreversibly obscuring critical tokens and causing severe error accumulation over long sequences. To address this limitation, we propose DLA, a dynamic memory modeling framework for multi-state linear attention. DLA introduces (i) Information-Aware Dynamic State Merging, which adaptively determines state boundaries based on token-level information variation, preserving high-resolution representations around semantic transitions while aggressively summarizing stable regions, and (ii) Capacity-Bounded Memory Modeling, which maintains a fixed-size, chronologically ordered state cache by selectively merging adjacent low-information states to control memory growth with minimal information loss. We pre-train DLA on two different linear attention models and evaluate on 16 datasets across three categories. Experimental results demonstrate the superiority of DLA over state-of-the-art.

0 Citations
0 Influential
19.5 Altmetric
97.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!