STAMP: 증거 기반 가이드 크레딧 할당을 통한 심층 검색 에이전트
STAMP: Provenance-Guided Credit Assignment for Deep Search Agents
심층 검색 에이전트를 위한 강화 학습은 주로 전체 경로 수준의 평가에 초점을 맞추고 있습니다. 여기에는 결과 정확도, 인용 정보를 고려한 보상 및 증거 범위가 포함됩니다. 그러나 관련 문서를 제시하는 행동에 대해서는 명확한 보상이 주어지지 않는데, 이를 우리는 '보상-크레딧 불일치'라고 부릅니다. 본 논문에서는 STAMP를 제안합니다. STAMP는 참조 기반 검증기를 사용하여 학습 시간 동안 생성된 증거 그래프에서 각 인용 문서가 엔터티 또는 관계를 지원하는지 판단하고, 첫 번째 노출 추적(first-exposure attribution)을 통해 지원되는 각 인용 정보를 해당 정보를 처음 제시한 행동으로 역추적합니다. 이러한 단계별 크레딧은 부호 유지 방식의 이점 조절(sign-preserving advantage modulation)을 통해 주입되며, 이를 통해 경로 수준 보상이나 그룹 내 트래젝토리의 상대적인 순위를 변경하지 않고도 이점을 재분배합니다. BrowseComp, BrowseComp-ZH 및 xbench-DS 데이터셋에서 STAMP는 동일한 SFT 초기화, 학습 데이터 및 검색 도구를 사용했을 때 GRPO 기준 모델보다 각각 +2.0/+5.5/+3.0 포인트의 성능 향상을 보였으며, 결과 기반 보상과 인용 정보 기반 보상 모두와 함께 사용할 수 있습니다. 구성 요소 분석을 통해 증거 기반 크레딧 신호와 부호 유지 방식의 이점 조절이 모두 성능 향상에 기여한다는 것을 확인했습니다.
Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the reward-credit mismatch. We propose STAMP, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and first-exposure attribution traces each supported citation back to the action that first surfaced it. This step credit is injected through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group. On BrowseComp, BrowseComp-ZH, and xbench-DS, STAMP improves the GRPO baseline by +2.0/+5.5/+3.0 points under matched SFT initialization, training data, and search tools, and composes with both outcome-only and citation-rubric base rewards. Component ablations confirm that the provenance-based credit signal and the sign-preserving advantage modulation each contribute to the gains.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.