2607.05798v1 Jul 07, 2026 cs.CV

질의 응답 전 분할: MLLM 시각적 추론을 위한 픽셀 기반 객체 지칭

Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning

Di Hu
Di Hu
Citations: 164
h-index: 6
Yake Wei
Yake Wei
Citations: 1,331
h-index: 13
Yuan Wang
Yuan Wang
Citations: 31
h-index: 2
Fengyun Rao
Fengyun Rao
Citations: 413
h-index: 5
Jing Lyu
Jing Lyu
Citations: 14
h-index: 2

최근 다중 모드 대규모 언어 모델(MLLM) 기술은 정적인 인식에서 상호 작용하는 시각-언어 추론으로 발전했으며, 이는 종종 '이미지를 사용한 사고'라고 불립니다. 이러한 추론 과정의 기본적인 작동 방식 중 하나는 관심 영역을 확대하여 더 자세한 시각 정보를 얻는 것입니다(이는 일반적으로 경계 상자로 표현됩니다). 본 논문에서는 질의 응답 전에 분할을 수행하는 방법인 SegAnswer를 제안합니다. 기존의 경계 상자 대신 픽셀 수준의 분할 마스크를 사용하여 관심 영역을 확대함으로써, 분할된 시각 정보는 더 정확한 관심 영역을 제공하고 불필요한 배경 및 방해 요소를 효과적으로 제거합니다. 또한, 분할된 시각 정보의 개별 패치들은 MLLM이 위치 임베딩을 통해 시각 토큰을 구조화하는 방식과 더욱 자연스럽게 연결됩니다. 실험에서는 고해상도 인식, 일반적인 인식, 환각 현상 등 다양한 벤치마크에서 SegAnswer를 평가했습니다. SegAnswer는 일관된 성능 향상을 보였으며, 분할 작업에서도 상당한 성능을 보여줌으로써 안정적인 픽셀 기반 객체 지칭 능력을 입증했습니다.

Original Abstract

Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with images''. A basic operation in this reasoning process is to zoom in on regions of interest (often represented with bounding boxes) to acquire finer visual details. In this paper, we propose \textbf{Seg}mentation before \textbf{Answer}ing (SegAnswer), which shifts the unit of zoom-in from the popular bounding box to pixel-level segmentation mask. By employing fine-grained masks to isolate the target area from cluttered environments, segmented visual input yields a more precise region of interest, effectively filtering out redundant background and interfering objects. Furthermore, the discrete patches of segmented visual input align more seamlessly with how MLLMs structure visual tokens via positional embeddings. In experiments, we evaluate SegAnswer across diverse benchmarks, including high-resolution perception, general perception, and hallucination. It achieves consistent improvements and also exhibits considerable performance on segmentation tasks, validating its capability for reliable pixel grounding.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!