어디를 보고 어떻게 판단할 것인가: 품질 인지 능력을 갖춘 주목 메커니즘을 활용한 해상도 불변 이미지 품질 평가
Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency
최근, 심층 및 다중 모드 모델 덕분에 참조 없는 이미지 품질 평가(NR IQA) 분야가 발전했지만, 여전히 많은 최첨단 시스템들이 기본적인 요구 사항 중 하나를 위반하고 있습니다. 이러한 시스템들은 원본 이미지의 중요한 품질 정보를 공격적인 리사이징을 통해 손실하거나, 다양한 해상도에 대한 일반화 성능이 부족하며, 불일치하는 MOS(Mean Opinion Score) 스케일을 가진 이질적인 IQA 데이터셋에 공동으로 학습할 수 없거나, 과도한 계산 비용이 필요합니다. 본 논문에서는 해상도에 영향을 받지 않고, 원본 해상도의 품질 정보를 보존하며, 다양한 주관적 연구로부터 학습하고, 계산 효율성이 높고 예산에 적응 가능한 모델인 **ReLIQS** ( **Re**solution-agnostic **L**earning for **I**mage **Q**uality with **S**aliency)을 제안합니다. ReLIQS는 CLIP 기반의 다중 스케일 패치 기반 아키텍처로, 이미지의 품질을 평가하기 위해 '어디를 주목해야 하는지'와 '어떻게 판단해야 하는지'를 학습합니다. 원본 해상도를 포함한 여러 해상도에서 고정 크기의 패치를 추출하고 CLIP 비전 백본을 사용하여 인코딩합니다. 경량화된 지각 중요도 추정기는 IQA에 특화된 중요도 맵을 예측하여 유용한 패치들을 선택하고, 잠재 품질 축 모듈은 이러한 패치의 임베딩을 단일 이미지 레벨 점수로 통합합니다. 다양한 해상도와 왜곡을 포함하는 실제 이미지, 합성 이미지 및 AIGC 벤치마크에서 ReLIQS는 동일하거나 감소된 계산 비용으로 CNN, CLIP 및 MLLM 기반의 강력한 기본 모델보다 더 나은 일반화 성능을 보입니다.
No-reference image quality assessment (NR IQA) has recently benefited from deep and multimodal models, yet many SOTA systems still violate at least one basic requirement: they either discard critical quality cues via aggressive resizing, fail to generalize across resolutions, cannot be jointly trained on heterogeneous IQA datasets with mismatched MOS scales, or require prohibitive computation. We present \textbf{ReLIQS}, a model for \textbf{Re}solution-agnostic \textbf{L}earning for \textbf{I}mage \textbf{Q}uality with \textbf{S}aliency, which is resolution-agnostic, preserves original-resolution quality cues, learns from multiple subjective studies, and remains computationally efficient and budget-adaptive. ReLIQS is a CLIP-based multiscale patch-driven architecture that learns both \emph{where to look} and \emph{how to judge} quality. Fixed-size patches are sampled across multiple resolutions, including the original resolution, and encoded with a CLIP vision backbone. A lightweight Perceptual Importance Estimator then predicts IQA-specific importance maps to select a small set of informative patches, and a Latent Quality Axis Module aggregates their embeddings into a single image-level score. Across authentic, synthetic, and AIGC benchmarks spanning diverse resolutions and distortions, ReLIQS generalizes better than strong CNN-, CLIP-, and MLLM-based baselines with matching or reduced computational cost.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.