2601.12294v1 Jan 18, 2026 cs.AI

ToolPRMBench: 도구 사용 에이전트를 위한 과정 보상 모델 평가 및 발전

ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents

Dawei Li
Dawei Li
Citations: 323
h-index: 4
Yuguang Yao
Yuguang Yao
Citations: 7
h-index: 2
Zhen Tan
Zhen Tan
Citations: 100
h-index: 5
Huan Liu
Huan Liu
Citations: 12
h-index: 2
Ruocheng Guo
Ruocheng Guo
Citations: 233
h-index: 5

보상 유도 탐색 방법은 복잡한 행동 공간에서의 샘플링과 탐색을 효과적으로 유도함으로써 도구 사용 에이전트의 성능을 향상시키는 데 강력한 잠재력을 입증했습니다. 이러한 탐색 방법의 핵심 설계는 과정 보상 모델(PRM)을 활용하여 단계별 보상을 제공함으로써 더 세밀한 모니터링을 가능하게 하는 것입니다. 그러나 도구 사용 환경에서 PRM을 위한 체계적이고 신뢰할 수 있는 평가 벤치마크는 부족한 실정입니다. 본 논문에서는 도구 사용 에이전트용 PRM을 평가하기 위해 특별히 설계된 대규모 벤치마크인 ToolPRMBench를 소개합니다. ToolPRMBench는 몇 가지 대표적인 도구 사용 벤치마크를 기반으로 구축되었으며, 에이전트의 궤적을 단계별 테스트 케이스로 변환합니다. 각 케이스에는 상호작용 기록, 정답 행동, 그럴듯하지만 틀린 대안, 그리고 관련 도구 메타데이터가 포함됩니다. 우리는 오프라인 샘플링을 활용하여 국소적인 단일 단계 오류를 분리하고, 온라인 샘플링을 활용하여 전체 에이전트 실행 과정에서 발생하는 현실적인 다단계 실패를 포착했습니다. 레이블 노이즈를 줄이고 데이터 품질을 보장하기 위해 다중 LLM 검증 파이프라인이 제안됩니다. 우리는 ToolPRMBench에서 대규모 언어 모델, 일반 PRM, 도구 특화 PRM을 대상으로 광범위한 실험을 수행했습니다. 실험 결과는 PRM의 효과성에 있어 명확한 차이를 보여주었으며, 도구 사용을 위한 특화된 PRM의 잠재력을 강조합니다. 코드와 데이터는 https://github.com/David-Li0406/ToolPRMBench 에 공개될 예정입니다.

Original Abstract

Reward-guided search methods have demonstrated strong potential in enhancing tool-using agents by effectively guiding sampling and exploration over complex action spaces. As a core design, those search methods utilize process reward models (PRMs) to provide step-level rewards, enabling more fine-grained monitoring. However, there is a lack of systematic and reliable evaluation benchmarks for PRMs in tool-using settings. In this paper, we introduce ToolPRMBench, a large-scale benchmark specifically designed to evaluate PRMs for tool-using agents. ToolPRMBench is built on top of several representative tool-using benchmarks and converts agent trajectories into step-level test cases. Each case contains the interaction history, a correct action, a plausible but incorrect alternative, and relevant tool metadata. We respectively utilize offline sampling to isolate local single-step errors and online sampling to capture realistic multi-step failures from full agent rollouts. A multi-LLM verification pipeline is proposed to reduce label noise and ensure data quality. We conduct extensive experiments across large language models, general PRMs, and tool-specialized PRMs on ToolPRMBench. The results reveal clear differences in PRM effectiveness and highlight the potential of specialized PRMs for tool-using. Code and data will be released at https://github.com/David-Li0406/ToolPRMBench.

7 Citations
0 Influential
33.486122886681 Altmetric
20.8 Score
Original PDF
8

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!