2608.02291v1 Aug 03, 2026 cs.AI

공통 접두사, 더 나은 신뢰도: 다중 에이전트 추론을 위한 적응형 라우팅

Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning

Hantao Yao
Hantao Yao
Citations: 11
h-index: 2
Wu Liu
Wu Liu
Citations: 22
h-index: 2
Yiqing Liu
Yiqing Liu
Citations: 18
h-index: 1
Yongdong Zhang
Yongdong Zhang
Citations: 64
h-index: 4
Zihao Wang
Zihao Wang
Citations: 0
h-index: 0

다중 에이전트 추론(MAR)은 반복적인 해결책 교환과 개선을 통해 추론의 신뢰성을 향상시킵니다. 기존의 적응형 MAR 방법들은 주로 쿼리 레벨의 레이블 또는 경로 레벨의 보상을 사용하여 라우팅 결정을 학습하지만, 이러한 거친 형태의 감독은 다단계 협업에서 개별 연산자의 상태에 따른 유용성을 정확하게 추정할 수 없습니다. 본 논문에서는 효율적인 적응형 MAR을 위한 공유 접두사 기반의 신뢰 할당 프레임워크인 TreeCredit을 제안합니다. TreeCredit의 핵심 아이디어는 경로 레벨의 결과를 직접 이전 결정에 귀속시키는 대신, 상태 일치 다운스트림 비교를 통해 연산자 유용성을 추정하는 것입니다. TreeCredit은 동일한 중간 상태에서 후보 연산자를 확장하여 공유 접두사 협업 트리를 구성하고, 각 상태-연산자 쌍에 대해 최종 정확성과 누적 추가 비용을 기반으로 정확도 우선 순위를 부여한 서피스 신뢰를 할당합니다. 이러한 구조화된 신뢰 값은 상태 지역 연산자 선호도로 변환되어 가벼운 쌍별 상태 라우터를 학습하는 데 사용되며, 이 라우터는 추론 과정에서 다음 허용 가능한 연산자를 동적으로 선택합니다. 여섯 가지 추론 벤치마크에 대한 실험 결과, TreeCredit은 정확도를 약간 향상시키면서 상당한 수준으로 추론 비용을 줄여 대표적인 MAR 방법보다 더 나은 정확도-비용 균형을 제공하는 것으로 나타났습니다.

Original Abstract

Multi-agent reasoning (MAR) improves reasoning reliability through iterative solution exchange and refinement. Existing adaptive MAR methods typically learn routing decisions from query-level labels or trajectory-level returns, but such coarse supervision cannot accurately estimate the state-conditioned utility of individual operators in multi-step collaboration. We propose TreeCredit, a shared-prefix credit assignment framework for efficient adaptive MAR. Its core insight is to estimate operator utility through state-matched downstream comparisons, rather than directly attributing trajectory-level outcomes to preceding decisions. TreeCredit constructs shared-prefix collaboration trees by expanding candidate operators from the same intermediate state and assigns each state--operator pair a correctness-prioritized suffix credit based on the terminal correctness and cumulative additional cost of its complete continuation. These structured credits are converted into state-local operator preferences to train a lightweight pairwise state router, which dynamically selects the next admissible operator during inference. Experiments on six reasoning benchmarks show that TreeCredit modestly improves accuracy while substantially reducing inference cost, achieving a better accuracy--cost trade-off than representative MAR methods.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!