문서 추출을 위한 비전-언어 모델의 신뢰도를 믿을 수 있을까요? ConfBench
Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction
지능형 문서 처리(IDP)에서 비전-언어 모델(VLM)은 자동화와 인간 검토 간의 추출 작업을 분배할 만큼 신뢰할 수 있는 신뢰도 점수를 필요로 합니다. 기존의 문서 벤치마크는 대부분 깨끗하고 품질이 높은 샘플로 구성되어 있어, 낮은 정확도 영역을 평가하기에는 데이터가 충분하지 않습니다. 본 연구에서는 핵심 정보 추출(KIE)을 위한 최초의 신뢰도 보정 특화 벤치마크인 ConfBench를 소개합니다. ConfBench는 다양한 문서 세트에 대해 20가지의 제어된 성능 저하 파이프라인을 적용하여 구축되었으며, 이를 통해 1,346개의 변형을 생성하고 전체 정확도 범위를 포괄하는 7만 개 이상의 엔티티 수준 평가를 수행했습니다. 본 연구에서는 네 가지 독점 모델과 세 가지 오픈 가중치 모델을 대상으로, 세 가지 입력 방식(모달리티)에서 언어화된 신뢰도 추정 방법과 로그 확률 방법을 사용하여 성능을 평가한 결과 다음과 같은 사실을 확인했습니다. (i) OCR+이미지 모달리티가 더 정확한 신뢰도 추정 결과를 제공합니다. (ii) 모델의 성능이 가장 중요한 요소이며, Claude 제품군 내에서는 신뢰도의 품질이 성능에 따라 단조롭게 증가하는 반면, 제품군 간 비교에서는 파라미터 수가 신뢰도를 잘 예측하지 못합니다. (iii) 모델별 신뢰도 보정 품질은 매우 다양하며, 거의 완벽한 수준부터 심각하게 과신되는 수준까지 나타납니다. 모델별 후처리 수정 방법을 통해 이러한 절대적인 신뢰도 값을 재조정하면 임계값 기반 라우팅에 영향을 주지 않으면서 순위 기반 운영 지표를 유지할 수 있습니다. (iv) 첫 번째 토큰을 집계하는 로그 확률 방법이 평균 토큰 집계 및 마진 집계를 지속적으로 능가합니다. 또한, 차별적인 성능 향상을 운영 비용 절감으로 변환하는 지표인 ECARB를 소개합니다. ConfBench는 신뢰도 추정기 및 신뢰도 보정 방법을 체계적으로 연구하여 신뢰할 수 있는 IDP 애플리케이션 배포를 가능하게 하기 위해 공개됩니다.
Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Existing document benchmarks are dominated by clean, high-quality samples, leaving low accuracy regions too sparse for calibration assessment. We introduce ConfBench, the first calibration-specific benchmark for key information extraction (KIE), built by applying 20 controlled degradation pipelines to a diverse document set, yielding 1,346 variants and 70K+ entity-level evaluations spanning the full accuracy spectrum. We evaluate four proprietary and three open-weight VLMs under verbalized and log-probability confidence estimation methods across three input modalities, and find: (i) OCR+Image modality results in more accurate confidence estimates; (ii) model capability is the dominant factor: within the Claude family confidence quality scales monotonically with capability, while across families parameter count is a poor predictor; (iii) calibration quality varies widely across models, from near-perfect to severely overconfident, and per-model post-hoc correction rescales these absolute confidence values for threshold-based routing without altering ranking-based operational metrics; and (iv) log-probability with first-token aggregation consistently outperforms mean-token and margin aggregations. We also introduce ECARB, a review-budget metric translating discriminative gains into operational savings. We release ConfBench publicly to enable systematic study of confidence estimators and calibration methods for trustworthy IDP application deployment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.