2606.26535v1 Jun 25, 2026 cs.CV

환각에서 실체 기반 이해로: CRISP를 활용한 시각 공간 지능 진단

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP

Yi Yu
Yi Yu
Citations: 9
h-index: 2
Zhixin Li
Zhixin Li
Citations: 31
h-index: 3

현재의 VLM(Visual Language Model) 평가 방식은 종종 언어적 편향을 실제적인 공간 추론 능력과 혼동하는 경향이 있습니다. 이를 해결하기 위해, 우리는 일관성을 통해 시각 공간 지능을 평가하는 새로운 구조적 진단 방법인 CRISP를 소개합니다. CRISP는 암묵적인 인지와 명시적인 추론 간의 정렬 관계를 분석하며, 기존의 블랙박스 질의응답 방식과는 달리, 3차원 장면 그래프와 오라클 개입 프로토콜을 사용하여 잠재적인 추론 능력을 인식적 제약으로부터 분리합니다. 이러한 세분화된 진단은 체계적인 인지-추론 단절 현상을 밝혀냅니다. 중요한 점은, 독점 모델들은 강력한 잠재적 추론 엔진을 가지고 있지만, 부정확한 측정값 추정과 암묵적인 구조적 표현을 활용하지 못하는 심각한 문제점을 안고 있다는 것입니다. 반면, 오픈 소스 모델들은 다단계 구성 추론 능력의 부족으로 인해 근본적인 한계를 보입니다. CRISP는 단순히 언어적 편향에 의존하여 '정확하게 답하기'에서 벗어나, '인지하고, 검증하고, 추론하는' 진정한 능력을 평가함으로써, 엔드-투-엔드 사후 훈련을 넘어선 다중 모달 정렬을 위한 엄격한 로드맵을 제시합니다. 코드와 데이터셋은 https://github.com/iiyamayuki/CRISP-Bench 에서 이용할 수 있습니다.

Original Abstract

Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning. Unlike traditional black-box QA, CRISP utilizes metric 3D Scene Graphs and an oracle intervention protocol to decouple latent reasoning capabilities from perceptual bottlenecks. This granular diagnosis uncovers a systematic perception-reasoning disconnect. Crucially, we reveal that while proprietary models possess robust latent reasoning engines, they suffer from inaccurate metric estimation and a critical failure to leverage their implicit structural representations. Conversely, open-source models remain fundamentally bottlenecked by their lack of multi-hop compositional reasoning. By shifting the focus from merely ``guessing correctly'' via language priors to genuinely ``perceiving, verifying, and reasoning,'' CRISP offers a rigorous roadmap for multimodal alignment beyond end-to-end post-training. The code and dataset are available at https://github.com/iiyamayuki/CRISP-Bench.

0 Citations
0 Influential
21.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!