2601.15703v1 Jan 22, 2026 cs.AI

에이전트 불확실성 정량화

Agentic Uncertainty Quantification

Jiaxin Zhang
Jiaxin Zhang
Citations: 57
h-index: 4
Prafulla Kumar Choubey
Prafulla Kumar Choubey
Texas A&M Univeristy
Citations: 931
h-index: 17
Kung-Hsiang Huang
Kung-Hsiang Huang
Citations: 204
h-index: 7
Caiming Xiong
Caiming Xiong
Citations: 520
h-index: 11
Chien-Sheng Wu
Chien-Sheng Wu
Citations: 40
h-index: 4

AI 에이전트가 장기 추론에서 인상적인 능력을 보여주었지만, 초기의 인식적 오류가 비가역적으로 전파되는 '할루시네이션의 나선(Spiral of Hallucination)'으로 인해 그 신뢰성은 심각하게 저해됩니다. 기존 방법들은 딜레마에 직면해 있습니다. 불확실성 정량화(UQ) 방법은 일반적으로 위험을 진단하기만 할 뿐 해결하지 못하는 수동적 센서 역할을 하는 반면, 자기 성찰 메커니즘은 지속적이거나 목적 없는 수정으로 인해 어려움을 겪습니다. 이러한 격차를 해소하기 위해, 우리는 언어화된 불확실성을 능동적인 양방향 제어 신호로 변환하는 통합된 '이중 프로세스 에이전트 UQ(AUQ)' 프레임워크를 제안합니다. 우리의 아키텍처는 두 가지 상호 보완적인 메커니즘으로 구성됩니다. 시스템 1(불확실성 인식 메모리, UAM)은 맹목적인 의사결정을 방지하기 위해 언어화된 신뢰도와 의미론적 설명을 암시적으로 전파하며, 시스템 2(불확실성 인식 성찰, UAR)는 이러한 설명을 합리적 단서로 활용하여 필요할 때만 타겟팅된 추론 시점 해결을 촉발합니다. 이를 통해 에이전트는 효율적인 실행과 깊은 숙고 사이의 균형을 동적으로 조절할 수 있습니다. 폐쇄 루프 벤치마크와 개방형 심층 연구 작업에 대한 광범위한 실험을 통해, 우리의 비학습(training-free) 접근 방식이 우수한 성능과 궤적 수준의 캘리브레이션을 달성함을 입증했습니다. 우리는 이 원칙에 입각한 AUQ 프레임워크가 신뢰할 수 있는 에이전트를 향한 중요한 진전이 될 것이라 믿습니다.

Original Abstract

Although AI agents have demonstrated impressive capabilities in long-horizon reasoning, their reliability is severely hampered by the ``Spiral of Hallucination,'' where early epistemic errors propagate irreversibly. Existing methods face a dilemma: uncertainty quantification (UQ) methods typically act as passive sensors, only diagnosing risks without addressing them, while self-reflection mechanisms suffer from continuous or aimless corrections. To bridge this gap, we propose a unified Dual-Process Agentic UQ (AUQ) framework that transforms verbalized uncertainty into active, bi-directional control signals. Our architecture comprises two complementary mechanisms: System 1 (Uncertainty-Aware Memory, UAM), which implicitly propagates verbalized confidence and semantic explanations to prevent blind decision-making; and System 2 (Uncertainty-Aware Reflection, UAR), which utilizes these explanations as rational cues to trigger targeted inference-time resolution only when necessary. This enables the agent to balance efficient execution and deep deliberation dynamically. Extensive experiments on closed-loop benchmarks and open-ended deep research tasks demonstrate that our training-free approach achieves superior performance and trajectory-level calibration. We believe this principled framework AUQ represents a significant step towards reliable agents.

12 Citations
1 Influential
8.5 Altmetric
56.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!