2605.28713v1 May 27, 2026 cs.AI

사고를 압축으로: 당신의 추론 모델은 은밀하게 컨텍스트 압축기입니다

Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor

Daiting Shi
Daiting Shi
Citations: 297
h-index: 7
Zhiyuan Sun
Zhiyuan Sun
Citations: 3
h-index: 1
Yuci Liang
Yuci Liang
Citations: 53
h-index: 3
Guoxin Ma
Guoxin Ma
Citations: 23
h-index: 2
Yibin Liu
Yibin Liu
Citations: 59
h-index: 4
Ke Chen
Ke Chen
Citations: 20
h-index: 3
Chengzhengxu Li
Chengzhengxu Li
Citations: 69
h-index: 5
Yan Wang
Yan Wang
Citations: 12
h-index: 2
Zhaohan Zhang
Zhaohan Zhang
Citations: 14
h-index: 2
Yue Zhang
Yue Zhang
Citations: 79
h-index: 3

컨텍스트 압축은 LLM 추론 속도 향상을 위해 긴 입력 컨텍스트를 최소한의 정보 손실로 단축하는 것을 목표로 합니다. 기존 방법들이 유망한 결과를 보여주었지만, 일반적으로 복잡한 압축 모듈에 의존하거나 압축 전용 학습을 필요로 하며, 이는 LLM의 고유한 능력을 충분히 활용하지 못합니다. 이에 반해, 본 연구는 사고 모델 자체가 관련 작업 정보 조직화를 통해 자연스럽게 긴 컨텍스트를 압축할 수 있음을 밝혀냅니다. 우리는 이러한 아이디어를 바탕으로, 사고 자체를 압축된 컨텍스트로 간주하는 새로운 압축 패러다임인 '사고를 압축 (Thinking as Compression, TaC)'을 제안합니다. TaC는 특정 전용 압축기를 사용하지 않고, 사고 모델에게 사고 과정을 생성하도록 유도하여 압축된 컨텍스트 역할을 수행하게 하며, 이는 대부분의 기존 압축 방법보다 우수한 성능을 보입니다. 또한, 원시적인 사고 결과물이 예산 제한 및 단축 경로 문제에 취약할 수 있다는 점을 고려하여, TaC-C (Thinking as Compression Constrained)를 제안합니다. 이는 간단한 보상 기반 최적화 프레임워크를 활용하여 LLM이 효율적이고 제어 가능한 압축된 컨텍스트 형태로 사고하도록 유도합니다. 네 가지 장문 컨텍스트 질의응답 벤치마크 실험 결과, TaC-C는 기존 방법들을 지속적으로 능가하는 성능을 보였습니다. 4배 및 8배의 압축 비율에서, TaC-C는 평균 F1 점수에서 각각 17.4%와 23.4%, 평균 정확 일치 점수(EM)에서 각각 15.7%와 21.7%로 가장 강력한 경쟁 모델을 능가했습니다.

Original Abstract

Context compression aims to shorten long context inputs with minimal information loss for LLM inference acceleration. While existing methods have shown promise, they typically rely on complex compression modules or compression-specific training, leaving the intrinsic capabilities of LLMs underexplored. In contrast, this work reveals that a thinking model itself can naturally compress long contexts by organizing task-relevant information. We thus derive Thinking as Compression (TaC), a new compression paradigm that treats thinking itself as compressed context. Without relying on specific dedicated compressor, TaC directly prompts the thinking model to generate thinking traces as the shortened context, already outperforming most representative compression methods. Further, given that raw thinking output may struggle with budget control and shortcut behaviors, we introduce Thinking as Compression Constrained (TaC-C), leveraging a simple reward-driven optimization framework to elicit intrinsic thinking as compact and controllable compressed context. Experiments across four long-context QA benchmarks demonstrate that TaC-C consistently outperforms existing baselines. At 4x and 8x compression ratios, it surpasses the strongest competitor by 17.4% and 23.4% in average F1, and by 15.7% and 21.7% in average Exact Match Score (EM), respectively.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!