2608.06128v1 Aug 06, 2026 cs.AI

검색 에이전트를 위한 문맥 정보 정책 최적화

Contextual Information Policy Optimization for Search Agents

Xingyu Guo
Xingyu Guo
Citations: 7
h-index: 1
Linlin Yang
Linlin Yang
Citations: 64
h-index: 5
Baochang Zhang
Baochang Zhang
Citations: 50
h-index: 5
Wei Chen
Wei Chen
Citations: 0
h-index: 0

검색 에이전트는 대규모 언어 모델의 정적인 파라미터 메모리 한계를 극복하고, 다단계 추론 과정에서 외부 증거를 획득하고 활용할 수 있도록 합니다. 복잡하거나 변화하는 정보를 다루는 지식 집약적 작업에서 검색 에이전트의 신뢰성은 관련 증거를 검색하는 것뿐만 아니라, 이를 사용하여 후속 추론을 안내하는 데에도 달려 있습니다. 그러나 기존 방법은 주로 최종 답변의 정확성 또는 중간 진행 상황에 대한 보상을 제공하며, 검색된 증거에 기반하여 수행되는 작업(post-retrieval actions)이 실제로 수행되었는지 직접적으로 평가하지 않습니다. 이러한 불일치는 사전 지식 중심의 추론을 유발합니다. 즉, 에이전트는 내부 지식을 바탕으로 결론을 내리고, 검색은 주로 이를 확인하는 데 사용되어 확증 편향과 비효율적인 증거 활용을 초래합니다. 이러한 문제를 해결하기 위해, 우리는 외부 증거 활용에 명시적으로 정책 최적화를 연계시키는 증거 중심 강화 학습 프레임워크인 문맥 정보 정책 최적화 (Contextual Information Policy Optimization, CIPO)를 제안합니다. CIPO는 검색된 정보의 영향을 받은 추론 작업에 대한 세분화된 보상을 제공하며, 동시에 전반적인 결과 보상과 결합하여 답변 정확성을 유지합니다. 이를 통해 CIPO는 증거와 무관한 추측을 억제하고, 검색된 사실이 후속 추론을 안내하거나 수정하는 데 도움이 되는 추론 경로를 장려합니다. 특히 CIPO는 인간의 과정에 대한 주석이나 추가적인 보상 모델이 필요하지 않습니다. 일곱 가지 도메인 내 및 도메인 외 벤치마크에서 수행한 광범위한 실험 결과, CIPO는 사전 지식 중심의 추론 빈도를 줄이고 대부분의 작업에서 뛰어난 성능을 달성함을 보여줍니다.

Original Abstract

Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For knowledge-intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant evidence but also on using it to guide subsequent reasoning. However, existing methods primarily reward final-answer correctness or intermediate progress, without directly assessing whether post-retrieval actions are grounded in the retrieved evidence. This misalignment encourages prior-driven reasoning: agents form conclusions based on internal knowledge and use retrieval mainly to confirm them, resulting in confirmation bias and inefficient evidence use. To address this issue, we propose Contextual Information Policy Optimization (CIPO), an evidence-oriented reinforcement learning framework that explicitly aligns policy optimization with external evidence use. CIPO assigns dense, turn-level credit to reasoning actions influenced by retrieved information, while combining this evidence-use signal with a global outcome reward to preserve answer correctness. With this manner, CIPO discourages evidence-detached guesses and promotes reasoning trajectories in which retrieved facts can guide or revise subsequent reasoning. Importantly, CIPO requires neither human process annotations nor an additional reward model. Extensive experiments on seven in-domain and out-of-domain benchmarks show that CIPO reduces the prevalence of prior-driven reasoning and achieves excellent performance on most tasks.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!