2608.03631v1 Aug 04, 2026 cs.CV

SEER: 자율적 근거 인터페이스를 활용한 제어된 공간 관계 분류

SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Xueqi Cheng
Xueqi Cheng
Citations: 1,916
h-index: 22
Qiang Qiu
Qiang Qiu
Citations: 115
h-index: 4

공간 관계 질문은 모델이 비교하기 전에 쿼리 대상 주체와 객체를 식별해야 합니다. 하지만 VLM(Vision-Language Model)은 두 개체를 모두 인식할 수 있지만, 잘못된 인스턴스를 선택하거나 모호한 전체적인 관점에서 답변하는 경우가 있습니다. 본 연구에서는 쿼리에 특화된 근거를 명시적으로 제시함으로써 이러한 문제를 완화할 수 있는지 질문하며, SEER(Self-grounded Evidence for Entity-Relation Reasoning)라는 학습이 필요 없는 추론 시점의 근거 인터페이스를 제안합니다. SEER는 쌍별 위치 정보 결정 과정에서 후보 관계를 숨기고, 쿼리에 특화된 관점을 구축하여 주체/객체의 역할을 명확하게 나타내며, 전체 이미지와 희소한 박스 정보를 상호 보완적인 증거로 활용합니다. 역으로 검증 가능한 관계 선택 프로토콜에서는, 선택적 개선 단계를 통해 엔티티 역할을 바꾸고, 정확히 하나의 시각적 상태가 해당 역 관계를 만족하는 경우에만 순방향 결정을 변경합니다. 모델 평가 전에 분리된 이미지 데이터셋인 GQA-Train900 테스트에서 SEER는 Full 모델 대비 +3.94 [2.17, 5.72]의 성능 향상을 보였으며, 레이블 독립적인 근거 순서 균형 조정 및 엔티티 이름이 고유한 535개 항목에서도 성능 향상이 유지되었습니다. 변경되지 않은 프로토콜은 세 가지 모델 모두에서 2,434개의 필터링된 EmbSpatial 쌍-관계 질문에 대해 +4.35에서 +11.79의 성능 향상을 보였습니다. 매칭 컨트롤 그룹을 통해 지역 집중과 역할 기반 조건을 분리했습니다. 이러한 결과는 쿼리에 특화된 근거 구축이 주요 개선 요소이며, 상호 일관성이 프로토콜별로 적용되는 작은 개선 사항임을 보여줍니다.

Original Abstract

Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the wrong instance or an ambiguous global view. We ask whether making query-specific evidence explicit can mitigate this failure and propose SEER (Self-grounded Evidence for Entity-Relation Reasoning), a training-free inference-time evidence interface for frozen VLMs. SEER hides candidate relations during pair localization, constructs a query-specific view with explicit subject/object roles, and retains the full image and sparse box geometry as complementary evidence. For relation-choice protocols with exact inverse support, an optional refinement swaps the entity roles and changes the forward decision only when exactly one visual state obeys the corresponding inverse relation. On an image-disjoint GQA-Train900 test frozen before model scoring, SEER pools to +3.94 [2.17,5.72] over Full; the gain remains positive under label-independent grounding-order counterbalancing and on the 535 rows whose entity names are unique. The unchanged protocol yields +4.35 to +11.79 on all 2,434 filtered EmbSpatial pair-relation questions across three models. Matched controls separate local refocus from role-explicit conditioning. These results establish query-specific evidence construction as the principal intervention, with reciprocal consistency as a smaller protocol-specific refinement.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!