2601.15778v1 Jan 22, 2026 cs.AI

에이전트 확신 보정 (Agentic Confidence Calibration)

Agentic Confidence Calibration

Jiaxin Zhang
Jiaxin Zhang
Citations: 57
h-index: 4
Caiming Xiong
Caiming Xiong
Citations: 520
h-index: 11
Chien-Sheng Wu
Chien-Sheng Wu
Citations: 40
h-index: 4

AI 에이전트는 수동적인 언어 모델에서 복잡하고 다단계의 작업을 수행하는 자율 시스템으로 빠르게 발전하고 있습니다. 그러나 실패 상황에서의 과잉 확신은 고위험 환경 배포에 있어 근본적인 장벽으로 남아 있습니다. 정적인 단일 턴 출력을 위해 구축된 기존의 보정 방법들은 궤적에 따른 오류 누적, 외부 도구로 인한 불확실성, 불투명한 실패 양상과 같은 에이전트 시스템의 고유한 문제들을 해결할 수 없습니다. 이러한 문제를 해결하기 위해, 우리는 최초로 '에이전트 확신 보정'이라는 문제를 도입하고, 에이전트의 전체 궤적에 걸쳐 거시적 역학에서 미시적 안정성에 이르는 풍부한 프로세스 수준 특징을 추출하는 새로운 진단 프레임워크인 '통합 궤적 보정(Holistic Trajectory Calibration, HTC)'을 제안합니다. 단순하고 해석 가능한 모델을 기반으로 하는 HTC는 8개의 벤치마크, 다수의 LLM, 그리고 다양한 에이전트 프레임워크 전반에서 보정 및 식별력 모두 강력한 베이스라인들을 일관되게 능가합니다. 성능을 넘어 HTC는 세 가지 필수적인 진전을 제공합니다. 실패 배후의 신호를 드러내어 해석 가능성을 제공하고, 재학습 없이 도메인 간 적용이 가능한 전이 가능성을 실현하며, 외부 도메인인 GAIA 벤치마크에서 최적의 보정(가장 낮은 ECE)을 달성하는 '일반 에이전트 보정기(GAC)'를 통해 일반화를 달성합니다. 종합하면, 이러한 기여들은 확신 보정을 위한 새로운 프로세스 중심의 패러다임을 확립하며, AI 에이전트의 신뢰성을 진단하고 향상시키기 위한 프레임워크를 제공합니다.

Original Abstract

AI agents are rapidly advancing from passive language models to autonomous systems executing complex, multi-step tasks. Yet their overconfidence in failure remains a fundamental barrier to deployment in high-stakes settings. Existing calibration methods, built for static single-turn outputs, cannot address the unique challenges of agentic systems, such as compounding errors along trajectories, uncertainty from external tools, and opaque failure modes. To address these challenges, we introduce, for the first time, the problem of Agentic Confidence Calibration and propose Holistic Trajectory Calibration (HTC), a novel diagnostic framework that extracts rich process-level features ranging from macro dynamics to micro stability across an agent's entire trajectory. Powered by a simple, interpretable model, HTC consistently surpasses strong baselines in both calibration and discrimination, across eight benchmarks, multiple LLMs, and diverse agent frameworks. Beyond performance, HTC delivers three essential advances: it provides interpretability by revealing the signals behind failure, enables transferability by applying across domains without retraining, and achieves generalization through a General Agent Calibrator (GAC) that achieves the best calibration (lowest ECE) on the out-of-domain GAIA benchmark. Together, these contributions establish a new process-centric paradigm for confidence calibration, providing a framework for diagnosing and enhancing the reliability of AI agents.

7 Citations
1 Influential
5.5 Altmetric
36.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!