UPAIR: 불확실성-진척도 정렬을 통한 추론 상태 진단 및 선택적 개입
UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention
테스트 시간 스케일링은 추가적인 추론 시간을 통해 대규모 추론 모델(LRM)의 문제 해결 능력을 향상시키지만, 과도한 사고와 부족한 사고를 악화시킬 수 있으며, 이는 우리가 '추론 상태-행동 불일치'라고 정의합니다. 이러한 불일치를 해결하려면 신뢰할 수 있는 추론 상태 진단이 필요하지만, 단일 신호 모니터는 모호한 증거를 제공하며, 제어 기반 방식은 종종 결과에 레이블이 붙은 지도 학습이나 모델 특정 보정에 의존합니다. 우리는 '불확실성-진척도 정렬 가설'을 제시하며, 이는 프록시 답변의 불확실성과 잠재적인 추론 진행률 간의 상대적 전환 시점이 건강 상태, 정체된 상태 및 준비 완료 상태를 구별하며, 이러한 상태는 서로 다른 후속 작업을 필요로 한다고 주장합니다. 이 통찰력을 바탕으로, 우리는 경량화된 불확실성 모니터링을 이벤트 기반의 공동 진단과 결합하고, 결과적인 상태를 모델의 기본 연속 작업, 선택적 전략 전환 또는 검증 지향적 중지로 매핑하는 학습이 필요 없는 프레임워크인 UPAIR를 제안합니다. 세 가지 LRM과 다섯 가지 교차 도메인 벤치마크에서, UPAIR의 정체 진단은 자연스러운 오류의 64.3%를 탐지하면서 동시에 정확한 샘플의 5.4%만을 오탐으로 판정하여 모델과 작업에 걸쳐 공유되는 동적인 추론 규칙성을 보여줍니다. 전체적으로, UPAIR는 정확도를 최대 16.67% 포인트 향상시키고 생성 토큰 수를 최대 29.64% 감소시켜 통합된 진단 및 개입의 효과를 입증하며, 온라인 진단의 비용은 자연스러운 생성 시간의 1% 미만입니다.
While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and underthinking, which we formulate as reasoning state--action mismatch. Resolving this mismatch requires reliable reasoning state diagnosis, yet single-signal monitors provide ambiguous evidence, while steering-based controllers often rely on outcome-labeled supervision or model-specific calibration. We introduce the Uncertainty--Progress Alignment Hypothesis, which posits that the relative transition timing of proxy answer uncertainty and latent reasoning progress distinguishes healthy, stagnant, and ready states that warrant different subsequent actions. Building on this insight, we propose UPAIR, a training-free framework that couples lightweight uncertainty monitoring with event-triggered joint diagnosis and maps the resulting state to native continuation, selective strategy switching, or verification-guided stopping. Across three LRMs and five cross-domain benchmarks, the stagnation diagnosis detects 64.3% of natural errors while flagging only 5.4% of correct samples, revealing a dynamic reasoning regularity shared across models and tasks. End to end, UPAIR improves accuracy by up to 16.67 percentage points and reduces generated tokens by up to 29.64%, demonstrating the effectiveness of its integrated diagnosis and intervention, while online diagnosis costs less than 1% of natural-generation time.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.