2606.13598v1 Jun 11, 2026 cs.AI

다중 에이전트 오케스트레이션의 보상 모델링

Reward Modeling for Multi-Agent Orchestration

Zixuan Ke
Zixuan Ke
Citations: 345
h-index: 9
Semih Yavuz
Semih Yavuz
Citations: 2,800
h-index: 25
Shafiq Joty
Shafiq Joty
Citations: 816
h-index: 18
Vishal Venkataramani
Vishal Venkataramani
Citations: 5
h-index: 2
Haizhou Shi
Haizhou Shi
Citations: 524
h-index: 6
King Yeung Tsang
King Yeung Tsang
Citations: 0
h-index: 0
Hao Wang
Hao Wang
Citations: 404
h-index: 4
Zihao Zhao
Zihao Zhao
Citations: 6
h-index: 2

대규모 언어 모델(LLM)을 기반으로 구축된 다중 에이전트 시스템(MAS)은 전문화된 에이전트를 조정하기 위한 효과적인 오케스트레이션을 필요로 하지만, 이러한 오케스트레이터의 학습은 제한된 감독과 높은 계산 비용 때문에 어려움을 겪습니다. 우리는 인간 어노테이션 없이 오케스트레이션 품질을 평가하는 자기 지도 방식의 프레임워크인 Orchestration Reward Modeling (OrchRM)을 제안합니다. OrchRM은 다중 에이전트 실행에서 생성된 중간 결과물을 활용하여 Bradley-Terry 보상 모델 학습을 위한 승/패 쌍을 구성합니다. 기존 MAS 테스트 시간 스케일링 및 오케스트레이터 학습 프레임워크와 달리, 비용이 많이 드는 하위 에이전트 시뮬레이션을 사용하지 않고 OrchRM은 직접 오케스트레이션 수준에서 작동하여 효율적이고 뛰어난 성능의 보상 기반 오케스트레이터 학습과 MAS 테스트 시간 스케일링을 가능하게 합니다. OrchRM은 토큰 사용량을 최대 10배까지 개선하고, 정확도 측면에서 MAS 테스트 시간 스케일링 성능을 최대 8%까지 향상시킵니다. 이러한 이점은 수학적 추론, 웹 기반 질의 응답 및 다중 단계 추론을 포함한 다양한 도메인에서 지속적으로 나타나며, 이는 오케스트레이션 수준의 보상 모델링이 강력한 다중 에이전트 오케스트레이션을 위한 확장 가능한 방향임을 보여줍니다. 관련 코드는 https://github.com/Wang-ML-Lab/OrchRM 에서 확인할 수 있습니다.

Original Abstract

Multi-Agent Systems (MAS) built on Large Language Models (LLMs) require effective orchestration to coordinate specialized agents, yet training such orchestrators is hindered by limited supervision and high computational cost. We propose Orchestration Reward Modeling (OrchRM), a self-supervised framework for evaluating orchestration quality without human annotations. OrchRM leverages intermediate artifacts from multi-agent executions to construct win-lose pairs for Bradley-Terry reward model training. Unlike existing MAS test-time scaling and orchestrator training frameworks that rely on costly sub-agent rollouts, OrchRM operates directly at the orchestration level, enabling efficient and high-performing reward-guided orchestrator training and MAS test-time scaling. OrchRM improves training efficiency by up to 10x in token usage while improving MAS test-time scaling performance by up to 8% in accuracy. These gains consistently transfer across multiple domains, including mathematical reasoning, web-based question answering, and multi-hop reasoning, demonstrating orchestration-level reward modeling as a scalable direction for robust multi-agent orchestration. Code will be available at https://github.com/Wang-ML-Lab/OrchRM.

1 Citations
0 Influential
32.5 Altmetric
6.9 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!