모델의 사전 지식이 시각적 증거와 충돌할 때: 선택적인 사전 보정(Prior Calibration)을 통한 상식 기반 환각 현상 완화
When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
이미지-텍스트 모델에서, 상식 기반 환각(Commonsense-Driven Hallucination, CDH)은 모델의 상식적인 사전 지식이 비정상적인 상태에 대한 명확한 시각적 증거를 압도할 때 발생합니다. 예를 들어, 뚜렷하게 여섯 개의 손가락을 가진 손이 보임에도 불구하고 모델이 그 손에 다섯 개의 손가락만 있다고 보고할 수 있습니다. 우리는 이러한 오류들이 체계적으로 나타나는 경향이 있음을 보여줍니다. 즉, 모델이 반사실적(Counterfactual, CF) 이미지에 대한 질문에 잘못 답하는 경우, 해당 답변은 종종 이미지를 보지 못했을 때 모델이 선호하는 후보와 일치합니다. 이 사전 지식을 무분별하게 억제하면 CF 오류를 수정할 수 있지만, 동시에 동일한 사전 지식이 도움이 되는 정상적인 상식(Commonsense, CS) 이미지에 대한 정확한 답변을 방해할 수도 있습니다. 따라서 우리는 선택적인 사전 보정(Selective Prior Calibration, SPC)이라는 방법을 제안합니다. SPC는 이미지 정보에 기반한 점수에 대해 후보 수준의 사전 선호도 추정치를 조정하며, 그 강도는 각 사례에 따라 다릅니다. 또한, 결과 점수 패턴이 다른 답변을 강력하게 지지할 때만 원래 예측을 수정합니다. 광범위한 실험 결과, SPC는 CF 이미지에 대한 정확도를 크게 향상시키면서 동시에 CS 이미지에 대한 정확도를 거의 그대로 유지하는 것을 보여줍니다. 더욱이, 이러한 이점은 CDH 유형, 후보 답변 순열 및 기타 충돌 벤치마크를 포함하여 다양한 조건에서 일반화됩니다. 반면, SPC는 이러한 충돌이 없는 벤치마크에서는 예측을 거의 변경하지 않습니다.
In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypical state. For example, a model may report that a visibly six-fingered hand has five fingers. We show that these errors are systematically directed: when a model answers a question about a counterfactual (CF) image incorrectly, its answer often coincides with the candidate it prefers without access to the image. Suppressing this prior indiscriminately can repair CF errors, but may also disrupt correct answers on matched commonsense (CS) images, where the same prior is helpful. We therefore propose Selective Prior Calibration (SPC), which subtracts candidate-level prior-preference estimates from image-conditioned scores with an instance-dependent strength and revises the original prediction only when the resulting score pattern strongly supports an alternative. Extensive experiments demonstrate that SPC substantially improves accuracy on CF images while largely preserving accuracy on matched CS images. Furthermore, these gains generalize across CDH categories, candidate-answer permutations, and other conflict benchmarks, while SPC rarely alters predictions on benchmarks without such conflicts.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.