2607.27919v1 Jul 30, 2026 cs.CL

대규모 메모리 디코더: 사전 학습된, 매개변수 기반의 장기 기억 모델

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Zhouhan Lin
Zhouhan Lin
Citations: 104
h-index: 5
Qipeng Guo
Qipeng Guo
Citations: 1,534
h-index: 10
Bowen Zhou
Bowen Zhou
Citations: 1,275
h-index: 8
Jiarui Wang
Jiarui Wang
Citations: 674
h-index: 15
Rubin Wei
Rubin Wei
Citations: 13
h-index: 2
Jiaqi Cao
Jiaqi Cao
Citations: 13
h-index: 2
Junming Zhang
Junming Zhang
Citations: 0
h-index: 0

디코더 전용 언어 모델은 장기 기억과 추론을 단일 매개 변수 집합으로 묶어, 메모리 용량을 독립적으로 확장하기 어렵게 만듭니다. Memory Decoder는 매개변수 기반의 장기 기억 모듈을 도입했지만, 상대적으로 작은 규모에서만 연구되었습니다. 본 논문에서는 Memory Decoder at Scale을 제시하며, 최대 69억 개의 파라미터를 갖는 메모리 모델을 개발하고 3000억 개의 토큰으로 사전 학습합니다. 이러한 데이터 규모에서는 표준 Faiss 파이프라인의 인덱싱 및 검색 비용이 매우 높아 비현실적입니다. 우리는 분산된 Faiss 인덱싱 및 검색 파이프라인과 함께, 희소하고 일괄적인 kNN 분포 로딩을 통해 이 병목 현상을 해결합니다. 모델 규모에 따라, 더 많은 매개 변수를 메모리에 할당하는 것이 기본 모델만 확장하는 것보다 더 나은 성능-매개변수 비율을 제공한다는 것을 확인했습니다. 17개의 벤치마크에서, 69억 개의 일반적인 메모리를 Pythia-410M과 결합하면 평균 점수가 29.86점에서 37.34점으로 향상되어, 총 매개변수 수가 39% 더 적은 Pythia-12B (37.24점)를 능가합니다. 0.6B에서 14B까지의 Qwen3 Base 모델에 대해, 1.7B 도메인 메모리는 세 가지 도메인 모두에서 평균 점수를 모든 규모에서 9점이 넘게 향상시킵니다. 전반적으로, 본 연구 결과는 사전 학습된 메모리를 독립적으로 확장하는 것이 언어 모델 성능을 향상시키는 더 효율적인 방법임을 보여줍니다.

Original Abstract

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!