2607.11172v1 Jul 13, 2026 cs.AI

STAMP: 증거 기반 가이드 크레딧 할당을 통한 심층 검색 에이전트

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

Yuchen Li
Yuchen Li
Citations: 71
h-index: 5
Xinran Chen
Xinran Chen
Citations: 740
h-index: 12
Ke Xu
Ke Xu
Citations: 0
h-index: 0
Han Xu
Han Xu
Citations: 0
h-index: 0
Yuqian Wang
Yuqian Wang
Citations: 43
h-index: 4
Zhixuan Li
Zhixuan Li
Citations: 0
h-index: 0
Xiaojia Liu
Xiaojia Liu
Citations: 0
h-index: 0
Changwo Wu
Changwo Wu
Citations: 0
h-index: 0
Jianqiang Xia
Jianqiang Xia
Citations: 0
h-index: 0

심층 검색 에이전트를 위한 강화 학습은 주로 전체 경로 수준의 평가에 초점을 맞추고 있습니다. 여기에는 결과 정확도, 인용 정보를 고려한 보상 및 증거 범위가 포함됩니다. 그러나 관련 문서를 제시하는 행동에 대해서는 명확한 보상이 주어지지 않는데, 이를 우리는 '보상-크레딧 불일치'라고 부릅니다. 본 논문에서는 STAMP를 제안합니다. STAMP는 참조 기반 검증기를 사용하여 학습 시간 동안 생성된 증거 그래프에서 각 인용 문서가 엔터티 또는 관계를 지원하는지 판단하고, 첫 번째 노출 추적(first-exposure attribution)을 통해 지원되는 각 인용 정보를 해당 정보를 처음 제시한 행동으로 역추적합니다. 이러한 단계별 크레딧은 부호 유지 방식의 이점 조절(sign-preserving advantage modulation)을 통해 주입되며, 이를 통해 경로 수준 보상이나 그룹 내 트래젝토리의 상대적인 순위를 변경하지 않고도 이점을 재분배합니다. BrowseComp, BrowseComp-ZH 및 xbench-DS 데이터셋에서 STAMP는 동일한 SFT 초기화, 학습 데이터 및 검색 도구를 사용했을 때 GRPO 기준 모델보다 각각 +2.0/+5.5/+3.0 포인트의 성능 향상을 보였으며, 결과 기반 보상과 인용 정보 기반 보상 모두와 함께 사용할 수 있습니다. 구성 요소 분석을 통해 증거 기반 크레딧 신호와 부호 유지 방식의 이점 조절이 모두 성능 향상에 기여한다는 것을 확인했습니다.

Original Abstract

Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the reward-credit mismatch. We propose STAMP, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and first-exposure attribution traces each supported citation back to the action that first surfaced it. This step credit is injected through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group. On BrowseComp, BrowseComp-ZH, and xbench-DS, STAMP improves the GRPO baseline by +2.0/+5.5/+3.0 points under matched SFT initialization, training data, and search tools, and composes with both outcome-only and citation-rubric base rewards. Component ablations confirm that the provenance-based credit signal and the sign-preserving advantage modulation each contribute to the gains.

1 Citations
0 Influential
6 Altmetric
31.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!