장기 탐색에서의 맥락 손상 진단 및 완화
Diagnosing and Mitigating Context Rot in Long-horizon Search
대규모 언어 모델(LLM)이 장기 과제에 점점 더 많이 활용됨에 따라, 방대한 문맥 정보가 일반적인 현상이 되었습니다. 증가된 문맥 길이가 모델의 성능을 저하시킨다는 '문맥 손상' 문제는 이러한 응용 분야에서 중요한 이슈로 부각되고 있습니다. 본 논문에서는 심층 탐색 시나리오에 초점을 맞춰, 문맥 손상 현상을 연구하고 이를 완화하는 전략을 제시합니다. 세 가지 벤치마크를 사용하여 네 가지 주요 오픈 소스 모델을 평가한 결과, 광범위하게 사용되지만 간과되었던 문맥 손상 현상이 확인되었습니다. 즉, 방대한 문맥은 모델이 직접적으로 포기하거나 불확실한 답변을 서둘러 제공하도록 유도하며, 이는 문맥의 길이가 증가함에 따라 더욱 심화됩니다. 가지치기 실험을 통해, 누적된 문맥과 문맥 손상 현상의 관계를 입증합니다. 또한, 문맥 관리 및 사후 거부 샘플링을 통해 이러한 문제를 완화하는 방법을 연구합니다. 문맥 관성의 경우, 성능, 비용 및 문맥 손상에 미치는 영향 측면에서 세 가지 범주에 걸쳐 7가지 다양한 방법을 체계적으로 평가하여 전략 선택 및 활용에 대한 명확한 지침을 제공합니다. 거부 샘플링의 경우, 문맥 손상을 고려하는 필터링 전략을 개발하고 세 가지 집계 방법에 따른 효과를 입증합니다. 마지막으로, 이 두 가지 접근 방식을 결합하여 성능을 더욱 향상시킬 수 있음을 보여줍니다.
Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon tasks. The concern that increasing context length degrades model capabilities, known as context rot, has become a central issue for these applications. In this paper, we focus on deep search scenarios, aiming to investigate the rot phenomenon and its mitigation strategies. By evaluating four flagship open-source models across three benchmarks, we reveal a prevalent but unnoticed rot phenomenon: extensive context causes models to directly give up or prematurely provide uncertain answers, and this issue is exacerbated as the context grows. Through pruning experiments, we demonstrate the relationship between the accumulated context and the rot phenomenon. Furthermore, we investigate mitigating this issue through context management and post-hoc rejection sampling. For context management, we systematically evaluate seven different methods across three categories, based on performance, cost, and impact on context rot, providing clear guidance for strategy selection and usage. For rejection sampling, we develop a rot-aware filtering strategy and demonstrate its effectiveness across three aggregation methods. Finally, we show that these two approaches can be combined for further performance improvements.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.