2608.02347v1 Aug 03, 2026 cs.AI

계층적 메모리를 활용한 맘바(Mamba): 장기 시퀀스 모델링에서의 표현 병목 현상 해결

Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling

Luziwei Leng
Luziwei Leng
Citations: 532
h-index: 12
Jianxiong Tang
Jianxiong Tang
Citations: 2
h-index: 1
Ruoyu Zhao
Ruoyu Zhao
Citations: 0
h-index: 0
Wei Zhang
Wei Zhang
Citations: 0
h-index: 0
Qinwen Wang
Qinwen Wang
Citations: 0
h-index: 0
Jieping Luo
Jieping Luo
Citations: 0
h-index: 0
Aoxiang Qin
Aoxiang Qin
Citations: 0
h-index: 0
Zhichao Lu
Zhichao Lu
Citations: 15
h-index: 2

맘바(Mamba)와 같은 순환 선형 어텐션 모델(RLA)은 트랜스포머의 대안으로서 효율적인 선형 시간 시퀀스 모델링을 제공하지만, 고정된 용량의 순환 상태는 장기 시퀀스 모델링에 제약을 가합니다. 인간 기억의 계층적 구조에서 영감을 받아, 본 연구에서는 이러한 한계를 극복하기 위해 계층적 메모리 맘바(HMM)를 제안합니다. 사전 학습된 맘바 모델을 기반으로, HMM은 경량화된 작업 메모리를 통합하여 백본의 은닉 상태에 내재된 빠른 감각 기억으로부터 느린 문단 수준 의미(PLS)를 추출합니다. 추출된 PLS는 이후 과제 관련 검색을 위해 지속적인 장기 기억으로 압축됩니다. 이러한 계층적 의미 정보 처리 방식은 RLA의 표현 병목 현상을 극복하며, 매개변수 학습을 통해 다른 장문 맥락 강화 맘바 모델에서 관찰되지 않는 교차 과제 일반화 능력을 부여합니다. Passkey Retrieval 및 LongBench-E 태스크에 대한 평가 결과, HMM은 강력한 맘바 기반 모델 대비 검색 성공률을 34.3~37.1% 향상시키고 추론 정확도를 1.6~14.2% 향상시키는 것을 보여주었으며, 이는 추가적인 매개변수 2% 증가와 미미한 학습 오버헤드만을 수반합니다.

Original Abstract

Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fixed-capacity recurrent states limit long-sequence modeling. Drawing inspiration from hierarchical human memory, we propose Hierarchical Memory Mamba (HMM) to address this limitation. Building upon a pre-trained Mamba backbone, HMM integrates a lightweight working memory that extracts slow paragraph-level semantics (PLS) from the fast sensory memory embedded in the backbone's hidden states. The PLS is subsequently compressed into persistent long-term memory for task-relevant retrieval. The hierarchical processing of semantic information overcomes the representation bottleneck of RLAs and endows HMM cross-task generalization through parametric learning, which is not observed in other long-context enhanced Mamba variants. Evaluations on Passkey Retrieval and LongBench-E tasks demonstrate that HMM improves retrieval success by 34.3--37.1% and reasoning accuracy by 1.6--14.2% over strong Mamba-based models, while adding only 2% extra parameters and with minimal training overhead.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!