2607.17486v1 Jul 20, 2026 cs.PF

SALT: 맥락 중요성을 고려한 어휘 트리 기반 장문 컨텍스트 압축

SALT: Salience-Aware Lexical Trie for Long-Context Compression

Shangqian Gao
Shangqian Gao
Citations: 13
h-index: 1
Oteo Mamo
Oteo Mamo
Citations: 0
h-index: 0
Hyunji Yi
Hyunji Yi
Citations: 0
h-index: 0
Weikuan Yu
Weikuan Yu
Citations: 11
h-index: 2
Joydhriti Choudhury
Joydhriti Choudhury
Citations: 0
h-index: 0

대규모 언어 모델(LLM)이 점점 더 긴 프롬프트를 처리함에 따라, 추론 시스템에서 계산 비용과 KV-캐시 메모리 사용량이 주요 병목 현상으로 부상했습니다. 기존의 입력 레벨 프롬프트 압축 방법은 이러한 문제를 해결하지만, 각 문장을 스칼라 관련성 점수로 순위를 매겨 문서를 비정형 단어 및 문장 풀로 취급합니다. 제한된 예산 하에서는, 문서의 주요 주제(들)가 예산을 소모하여 덜 빈번하지만 작업과 관련된 주제들이 버려지는 '주제 붕괴' 현상이 발생합니다. 따라서, 주제적 범위를 유지하려면 예산을 개별 문장에 대한 점수 매기기가 아닌 반복되는 주제에 할당해야 합니다. 이를 위해, 우리는 모델에 독립적인 추출 프레임워크인 SALT를 제안합니다. SALT는 각 문장의 키워드를 문서의 주제 구조를 나타내는 가벼운, 재사용 가능한 '프록시'인 문장 빈도(SF) 순으로 정렬된 트리에 구성합니다. 이 트리 기반 구성은 메모리 할당을 최적화하고 주요 주제가 예산을 독점하는 것을 방지합니다. 멀티-앵커 검색은 쿼리 키워드로 레이블이 지정된 모든 깊이의 트리 노드를 활성화하며, 트리는 대화 턴 전체에서 유지되어 문서를 재인코딩하지 않고 다중 턴 사용을 지원합니다. SALT는 문서의 주제를 보존함으로써 장문 컨텍스트 프롬프트의 사전 채우기 계산 및 메모리 비용을 줄이는 동시에 디코딩 시간을 목표로 하는 KV-캐시 방법과 함께 사용할 수 있습니다.

Original Abstract

As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottlenecks in inference systems. Existing input-level prompt compression methods address this, but rank each sentence by a scalar relevance score, treating the document as an unstructured pool of words and sentences. Under tight budgets, this causes theme collapse, where the dominant theme(s) of a document consumes the budget, discarding less-frequent yet task-relevant themes. Preserving thematic coverage instead requires allocating the budget across recurring themes rather than scoring sentences in isolation. To this end, we propose SALT, a model-agnostic extractive framework that organizes per-sentence keywords into a trie ordered by sentence frequency (SF), a lightweight, reusable proxy for document thematic structure. This trie-based organization smooths memory allocation and prevents dominant themes from monopolizing the budget. Multi-anchor retrieval activates trie nodes labeled by query keywords at any depth, and the trie persists across dialogue turns, supporting multi-turn use without re-encoding the document. By preserving document themes, SALT reduces the prefill computation and memory cost of long-context prompts while remaining composable with KV-cache methods that target decoding-time latency and memory.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!