경로 기반 보상을 넘어: 그래프 모델링을 통한 주체적 탐색에서의 단계별 기여도 할당
Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling
주체적 탐색(Agentic Search)에서, 경로 수준의 결과 보상은 개별 단계들의 행동 기여도를 정확하게 반영하지 못하며, 기존의 단계별 보상 방법은 일반적으로 비용이 많이 드는 트리 샘플링에 의존합니다. 본 연구에서는 세계 지식을 잠재적인 세계 그래프로 보고, 각 IS(Interactive Search) 작업을 잠재된 작업 그래프 내에서의 탐색으로 간주합니다. 효과적인 단계는 답변 노드 방향으로 그래프의 진행을 가져와야 합니다. 이러한 사전 지식을 바탕으로, 우리는 훈련 시점의 개체-관계(ER) 그래프에서 답변 노드까지의 거리를 기준으로 새로 검색되고 인용된 개체의 기여도를 평가하는 단계별 프로세스 보상인 Graph-Distance Contribution Reward (GDCR)를 제안합니다. 또한, GDCR을 단계별 이점으로 변환하고, 이를 경로 수준의 결과 이점과 결합하는 Step Advantage Policy Optimization (SAPO) 방법을 제시합니다. 네 가지 어려운 벤치마크에서의 실험은 본 연구 방법의 효과성을 입증합니다.
In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward methods typically rely on costly tree sampling. We view world knowledge as a latent world graph and each IS task as search within a latent task graph, where effective steps should make graph progress toward the answer node. Based on this prior, we propose Graph-Distance Contribution Reward (GDCR), a step-level process reward that scores newly-retrieved and newly-cited entities by their distance to the answer node in a training-time Entity-Relation (ER) graph. We further propose Step Advantage Policy Optimization (SAPO), which converts GDCR into step-level advantages and combines them with trajectory-level outcome advantages. Experiments on four challenging benchmarks validate the effectiveness of our method.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.