2604.17950v1 Apr 20, 2026 cs.AI

CADMAS-CTX: 맥락 기반 능력을 활용한 다중 에이전트 위임

CADMAS-CTX: Contextual Capability Calibration for Multi-Agent Delegation

Chuhang Qiao
Chuhang Qiao
Citations: 0
h-index: 0

본 연구는 다중 에이전트 위임 문제를 다룰 때, 에이전트의 능력이 단순히 숙련도 수준으로 고정되는 것이 아니라, 작업 맥락에 따라 달라진다는 보다 강력하고 현실적인 가정을 기반으로 재검토합니다. 코딩 에이전트는 짧은 독립적인 수정 작업에서는 뛰어날 수 있지만, 장기적인 디버깅 작업에서는 실패할 수 있으며, 플래너는 얕은 작업에서는 잘 수행되지만, 복잡한 의존성 관계가 있는 작업에서는 성능이 저하될 수 있습니다. 따라서 정적인 숙련도 기반 능력 프로필은 다양한 상황을 평균화하기 때문에 체계적인 위임 오류를 초래할 수 있습니다. 본 연구에서는 맥락 기반 능력 교정 프레임워크인 CADMAS-CTX를 제안합니다. CADMAS-CTX는 각 에이전트, 숙련도, 그리고 거친 맥락 범주에 대해, 해당 작업 공간 영역에서의 안정적인 경험을 나타내는 베타 사후 분포를 유지합니다. 위임은 사후 평균과 불확실성 페널티를 결합한 위험 감지 점수를 사용하여 수행됩니다. 이를 통해 에이전트는 동료가 더 나은 성능을 보이는 경우에만 위임하며, 이러한 판단이 충분한 증거에 의해 뒷받침되는 경우에만 위임합니다. 본 논문은 세 가지 기여를 합니다. 첫째, 계층적 맥락 기반 능력 프로필을 통해 정적인 숙련도 기반 신뢰도를 맥락에 따라 달라지는 사후 분포로 대체합니다. 둘째, 맥락 기반 강도학 이론에 기반하여, 충분한 맥락 다양성이 존재하는 경우, 맥락 인지 라우팅이 정적 라우팅보다 누적 후회를 줄인다는 것을 형식적으로 증명하고, 편향-분산 트레이드오프를 명확히 합니다. 셋째, GAIA 및 SWE-bench 벤치마크를 사용하여 제안하는 방법을 경험적으로 검증합니다. GPT-4o 에이전트를 사용하는 GAIA 환경에서 CADMAS-CTX는 0.442의 정확도를 달성하여, 정적 기준 성능인 0.381과 AutoGen의 성능인 0.354를 능가하며, 95% 신뢰 구간에서 겹치지 않습니다. SWE-bench Lite 환경에서는 해결률을 22.3%에서 31.4%로 향상시켰습니다. 추가 분석 결과, 불확실성 페널티는 맥락 태깅 오류에 대한 강건성을 향상시키는 것으로 나타났습니다. 본 연구의 결과는 맥락 기반 교정과 위험 감지 위임이 정적인 전역 숙련도 할당보다 다중 에이전트 협업을 크게 향상시킬 수 있음을 보여줍니다.

Original Abstract

We revisit multi-agent delegation under a stronger and more realistic assumption: an agent's capability is not fixed at the skill level, but depends on task context. A coding agent may excel at short standalone edits yet fail on long-horizon debugging; a planner may perform well on shallow tasks yet degrade on chained dependencies. Static skill-level capability profiles therefore average over heterogeneous situations and can induce systematic misdelegation. We propose CADMAS-CTX, a framework for contextual capability calibration. For each agent, skill, and coarse context bucket, CADMAS-CTX maintains a Beta posterior that captures stable experience in that part of the task space. Delegation is then made by a risk-aware score that combines the posterior mean with an uncertainty penalty, so that agents delegate only when a peer appears better and that assessment is sufficiently well supported by evidence. This paper makes three contributions. First, a hierarchical contextual capability profile replaces static skill-level confidence with context-conditioned posteriors. Second, based on contextual bandit theory, we formally prove context-aware routing achieves lower cumulative regret than static routing under sufficient context heterogeneity, formalizing the bias-variance tradeoff. Third, we empirically validate our method on GAIA and SWE-bench benchmarks. On GAIA with GPT-4o agents, CADMAS-CTX achieves 0.442 accuracy, outperforming static baseline 0.381 and AutoGen 0.354 with non-overlapping 95% confidence intervals. On SWE-bench Lite, it improves resolve rate from 22.3% to 31.4%. Ablations show the uncertainty penalty improves robustness against context tagging noise. Our results demonstrate contextual calibration and risk-aware delegation significantly improve multi-agent teamwork compared with static global skill assignments.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!