2608.01792v1 Aug 03, 2026 cs.AI

문서 추출을 위한 비전-언어 모델의 신뢰도를 믿을 수 있을까요? ConfBench

Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction

Md Mofijul Islam
Md Mofijul Islam
Citations: 7
h-index: 1
Spencer Romo
Spencer Romo
Citations: 3
h-index: 1
Bob Strahan
Bob Strahan
Citations: 2
h-index: 1
Boyi Xie
Boyi Xie
Citations: 2,596
h-index: 7
Diego Socolinsky
Diego Socolinsky
Citations: 261
h-index: 3
Mohammad Rostami
Mohammad Rostami
Citations: 28
h-index: 4
P. Roy
P. Roy
Citations: 1
h-index: 1
Sujitha Martin
Sujitha Martin
Citations: 1,412
h-index: 20
Renhao Xue
Renhao Xue
Citations: 0
h-index: 0

지능형 문서 처리(IDP)에서 비전-언어 모델(VLM)은 자동화와 인간 검토 간의 추출 작업을 분배할 만큼 신뢰할 수 있는 신뢰도 점수를 필요로 합니다. 기존의 문서 벤치마크는 대부분 깨끗하고 품질이 높은 샘플로 구성되어 있어, 낮은 정확도 영역을 평가하기에는 데이터가 충분하지 않습니다. 본 연구에서는 핵심 정보 추출(KIE)을 위한 최초의 신뢰도 보정 특화 벤치마크인 ConfBench를 소개합니다. ConfBench는 다양한 문서 세트에 대해 20가지의 제어된 성능 저하 파이프라인을 적용하여 구축되었으며, 이를 통해 1,346개의 변형을 생성하고 전체 정확도 범위를 포괄하는 7만 개 이상의 엔티티 수준 평가를 수행했습니다. 본 연구에서는 네 가지 독점 모델과 세 가지 오픈 가중치 모델을 대상으로, 세 가지 입력 방식(모달리티)에서 언어화된 신뢰도 추정 방법과 로그 확률 방법을 사용하여 성능을 평가한 결과 다음과 같은 사실을 확인했습니다. (i) OCR+이미지 모달리티가 더 정확한 신뢰도 추정 결과를 제공합니다. (ii) 모델의 성능이 가장 중요한 요소이며, Claude 제품군 내에서는 신뢰도의 품질이 성능에 따라 단조롭게 증가하는 반면, 제품군 간 비교에서는 파라미터 수가 신뢰도를 잘 예측하지 못합니다. (iii) 모델별 신뢰도 보정 품질은 매우 다양하며, 거의 완벽한 수준부터 심각하게 과신되는 수준까지 나타납니다. 모델별 후처리 수정 방법을 통해 이러한 절대적인 신뢰도 값을 재조정하면 임계값 기반 라우팅에 영향을 주지 않으면서 순위 기반 운영 지표를 유지할 수 있습니다. (iv) 첫 번째 토큰을 집계하는 로그 확률 방법이 평균 토큰 집계 및 마진 집계를 지속적으로 능가합니다. 또한, 차별적인 성능 향상을 운영 비용 절감으로 변환하는 지표인 ECARB를 소개합니다. ConfBench는 신뢰도 추정기 및 신뢰도 보정 방법을 체계적으로 연구하여 신뢰할 수 있는 IDP 애플리케이션 배포를 가능하게 하기 위해 공개됩니다.

Original Abstract

Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Existing document benchmarks are dominated by clean, high-quality samples, leaving low accuracy regions too sparse for calibration assessment. We introduce ConfBench, the first calibration-specific benchmark for key information extraction (KIE), built by applying 20 controlled degradation pipelines to a diverse document set, yielding 1,346 variants and 70K+ entity-level evaluations spanning the full accuracy spectrum. We evaluate four proprietary and three open-weight VLMs under verbalized and log-probability confidence estimation methods across three input modalities, and find: (i) OCR+Image modality results in more accurate confidence estimates; (ii) model capability is the dominant factor: within the Claude family confidence quality scales monotonically with capability, while across families parameter count is a poor predictor; (iii) calibration quality varies widely across models, from near-perfect to severely overconfident, and per-model post-hoc correction rescales these absolute confidence values for threshold-based routing without altering ranking-based operational metrics; and (iv) log-probability with first-token aggregation consistently outperforms mean-token and margin aggregations. We also introduce ECARB, a review-budget metric translating discriminative gains into operational savings. We release ConfBench publicly to enable systematic study of confidence estimators and calibration methods for trustworthy IDP application deployment.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!