2607.12375v1 Jul 14, 2026 cs.CV

IQA-T1: 도구 기반 시각적 증거 추론을 통한 이미지 품질 평가

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

Qifeng Chen
Qifeng Chen
Citations: 85
h-index: 4
Botong Geng
Botong Geng
Citations: 0
h-index: 0
Lei Zhang
Lei Zhang
Citations: 6
h-index: 1
Jinjian Wu
Jinjian Wu
Citations: 114
h-index: 7
Jiaqi Tang
Jiaqi Tang
Citations: 62
h-index: 3
Wei Wei
Wei Wei
Citations: 8
h-index: 2
Yingying Yan
Yingying Yan
Citations: 0
h-index: 0
Jianmin Chen
Jianmin Chen
Citations: 3,258
h-index: 3

개방형 환경에서의 이미지 품질 평가(IQA)는 일반화 능력과 해석 가능성의 제한으로 인해 여전히 어려운 과제입니다. 최근 다중 모드 대규모 언어 모델(MLLM)에 기반한 접근 방식은 품질 예측을 위한 텍스트 추론을 도입했지만, 이러한 방법은 의미적으로 편향된 내부 표현에 크게 의존하여 저수준의 인지적 왜곡에 민감하지 않습니다. 본 논문에서는 MLLM 추론에 명시적인 인지적 관찰을 추가하는 도구 기반 시각적 증거 추론 프레임워크인 IQA-T1을 제안합니다. 추론 과정에서 모델은 자동으로 특수 분석 도구를 호출하여 노이즈 잔류 맵, 그래디언트 통계 및 주파수 스펙트럼과 같은 구조화된 시각적 증거를 생성하며, 이러한 증거는 점진적으로 추론 과정에 통합됩니다. 이러한 패러다임을 지원하기 위해, 도구에서 생성된 증거를 기반으로 하는 11,000개의 다중 모드 추론 체인을 포함하는 데이터셋인 Q-Tool을 구축했습니다. 일곱 개의 IQA 벤치마크에 대한 광범위한 실험 결과, IQA-T1은 전반적으로 가장 뛰어난 성능을 보이며 해석 가능하고 증거 기반의 품질 평가를 제공합니다. 코드 및 데이터셋은 https://github.com/zibuyu-02/IQA-T1 에서 확인할 수 있습니다.

Original Abstract

Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches based on multimodal large language models (MLLMs) introduce textual reasoning for quality prediction, yet their judgments rely heavily on semantically biased internal representations, making them insensitive to low-level perceptual degradations. We propose IQA-T1, a tool-based visual evidence reasoning framework that augments MLLM reasoning with explicit perceptual observations. During inference, the model autonomously invokes specialized analysis tools to generate structured visual evidence, such as noise residual maps, gradient statistics, and frequency spectra, which are progressively integrated into the reasoning process. To support this paradigm, we construct Q-Tool, a dataset containing 11k multimodal reasoning chains grounded in tool-generated evidence. Extensive experiments on seven IQA benchmarks show that IQA-T1 achieves the best overall performance across datasets while producing interpretable and evidence-grounded quality assessments. Code and dataset are available at https://github.com/zibuyu-02/IQA-T1.

0 Citations
0 Influential
23.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!