Tool-IQA: 간단한 도구를 활용하여 이미지 품질 평가 성능 향상
Tool-IQA: Augmenting Image Quality Assessment with Simple Tools
최근 비전-언어 모델(VLMs)은 이미지 품질 평가(IQA) 분야에서 점점 더 많이 활용되고 있습니다. 그러나 현재의 방법들은 대부분 정적인 단일 스코어링 방식을 사용하는데, 이는 인간이 세부 사항 및 미세한 왜곡을 확인하기 위해 시야를 선택적으로 조정하는 등 동적인 시각 검사를 통해 이미지 품질을 평가한다는 사실과는 다릅니다. 특히, 단일 관찰에만 의존하면 두 가지 주요 한계가 발생합니다. 첫째, 전체적인 규모에서 이미지를 인식하는 것만으로는 미세한 지역적 세부 사항을 평가하기 어렵습니다. 둘째, 이미지의 원래 강도 분포가 시야를 압도하여 이미지 품질 검사에 충분하지 않을 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 수동 스코어링 방식에서 벗어나 도구를 활용하는 워크플로우로 평가 방식을 전환하는 Tool-IQA를 제안합니다. 특히, 우리는 VLM에 간단하지만 효과적인 시야 조정 도구를 제공합니다: 지역적 세부 사항을 검사할 수 있는 확대경(Magnifier)과, 가시성을 향상시키고 숨겨진 왜곡을 드러낼 수 있는 감마 보정기(Gamma Corrector). 평가는 초기 관찰 및 참고 사항 작성, 도구를 활용한 심층적인 검사, 그리고 최종적으로 보정된 품질 점수를 산출하는 구조화된 파이프라인으로 구성됩니다. 또한, 효율적이고 목적 지향적인 도구 사용을 위해, 단순히 도구 사용을 장려하는 것이 아니라 긍정적인 기여를 할 수 있는 도구 상호 작용에 대해 보상을 제공하는 배치 인식 학습 전략을 도입했습니다. 다양한 IQA 벤치마크에서의 실험 결과는, 효과적인 도구 활용과 보정된 평가를 통해 제안하는 Tool-IQA가 기존의 최첨단 모델보다 훨씬 우수한 성능을 발휘한다는 것을 보여줍니다. 예를 들어, 어려운 CLIVE 데이터셋에서 0.854의 PLCC 값을 달성했습니다.
Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA). However, current methods typically employ a static one-shot scoring paradigm, despite the fact that humans assess image quality through dynamic visual inspection, e.g., selectively adjusting views to verify details and subtle artifacts. Specifically, relying solely on a single-pass observation introduces two primary limitations: first, perceiving the image only at a global scale restricts the assessment of finer local details; second, the original intensity distribution of the image may overwhelm the visibility, leading to insufficient inspection of image quality. To address these issues, we propose Tool-IQA, shifting the assessment mechanism from passive scoring to a tool-augmented workflow. In particular, we equip VLMs with simple yet effective view tools: a Magnifier to inspect local details, and a Gamma Corrector to uncover visibility and hidden artifacts. The assessment follows a structured pipeline that consists of an initial observation with rubric notes, a tool-augmented in-depth inspection, and a final quantification for calibrated quality score. Furthermore, to ensure efficient and purposeful tool callings, we introduce a batch-aware training strategy to reward tool interactions that can yield positive contributions rather than simply encouraging usage. Experiments on a variety of IQA benchmarks demonstrate that, with effective tool calling and calibrated assessment, our proposed Tool-IQA significantly outperforms existing state-of-the-art models, e.g., it achieves a PLCC of 0.854 on the challenging CLIVE dataset.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.