2606.05784v1 Jun 04, 2026 cs.AI

TAPO: 도구 인지 정책 최적화 - 다중 모드 검색 에이전트를 위한 신용 이전 기법

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents

Guojun Yin
Guojun Yin
Citations: 265
h-index: 8
Hang He
Hang He
Citations: 67
h-index: 5
Xiaohan Wang
Xiaohan Wang
Citations: 101
h-index: 6
Jiajun Chai
Jiajun Chai
Citations: 116
h-index: 6
Chuhuai Yue
Chuhuai Yue
Citations: 50
h-index: 3
Fenghe Tang
Fenghe Tang
University of Science and Technology of China
Citations: 565
h-index: 10
S. K. Zhou
S. K. Zhou
Citations: 57
h-index: 5
Chengqi Dong
Chengqi Dong
Citations: 41
h-index: 3
Yandong Liu
Yandong Liu
Citations: 3
h-index: 1

본 논문에서는, 도구를 활용하는 다중 모드 검색 에이전트에서 GRPO(Generalized Reinforcement Policy Optimization)의 주요 문제점인 '신용 오배분'을 체계적으로 분석하고 공식화합니다. GRPO는 궤적 레벨의 이점을 모든 토큰에 균일하게 분산함으로써, 실패한 궤적 내에서도 유용한 도구 사용 단계를 무가치한 단계와 동일하게 벌점으로 처리하는 문제를 야기합니다. 또한, 우리는 이러한 현상의 규모를 실증적으로 정량화했습니다. 분석 결과, 절반 이상의 실패한 궤적과 실패한 도구 사용 동작에서 수정 가능한 신용 오배분이 발생하는 것을 확인했으며, 이는 상당한 양의 학습 신호가 낭비되고 있으며, 구조적으로 활용될 수 있음을 보여줍니다. 이러한 통찰력을 바탕으로, 우리는 정보 획득 도구의 매개변수 결정론적 특성을 활용하는 '도구 인지 정책 최적화(TAPO)' 기법을 제안합니다. TAPO는 현재 학습 배치 내에서 반사실적 증거를 구성하고, 신뢰도 기반의 보수적인 이점 교정 메커니즘을 통해 잘못 할당된 부정적인 신용을 보완합니다. TAPO는 추가적인 어노테이션, 모델 또는 샘플링 과정이 필요 없으며, 계산 오버헤드가 미미합니다. 다양한 다중 모드 검색 벤치마크에서 TAPO는 GRPO, GSPO 및 SAPO와 같은 주요 강화 학습 알고리즘에 대해 일관되고 간편하게 성능 향상을 제공합니다. 본 논문의 코드와 모델은 채택 결정 후 공개될 예정입니다.

Original Abstract

We identify and formally characterize credit misassignment as a systematic failure mode of GRPO in tool-augmented multimodal search agents: its uniform broadcast of trajectory-level advantages to all tokens causes valuable tool-use steps in failing trajectories to be penalized no differently from valueless ones. We further empirically quantify the scale of this phenomenon. Over half of failing trajectories and failing tool-use actions exhibit correctable credit misassignment, demonstrating that the wasted training signal is both substantial and structurally exploitable. Building on this insight, we propose Tool-Aware Policy Optimization (TAPO), which exploits the parameter-determinism property of information-acquisition tools: similar call parameters define equivalent information-acquisition actions and should therefore share comparable action credit. TAPO constructs counterfactual witnesses within the current training batch and compensates misassigned negative credit via confidence-gated conservative advantage correction. It requires no additional annotation, models, or sampling, and introduces negligible computational overhead. Across multiple multimodal search benchmarks, TAPO delivers consistent, plug-and-play improvements over strong baselines for three mainstream RL algorithms (GRPO, GSPO, and SAPO). Our code and models will be publicly released upon acceptance.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!