2607.27145v1 Jul 29, 2026 cs.CV

의사 결정에 중요한 응용 분야를 위한 다중 모드 대규모 언어 모델에서 설명 가능하고 효율적인 공간 추론

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

Rajarshi Roy
Rajarshi Roy
Citations: 10
h-index: 2
Piyush Jain
Piyush Jain
Citations: 0
h-index: 0
Kousik Dasgupta
Kousik Dasgupta
Citations: 3
h-index: 1
Subarna Tripathi
Subarna Tripathi
Citations: 2
h-index: 1

다중 모드 대규모 언어 모델(MLLM)이 로봇 공학, 에이전트 인공지능 및 안전 감시와 같은 의사 결정에 중요한 시스템에 점점 더 많이 사용됨에 따라, MLLM의 공간적 판단의 불투명성은 운영자의 신뢰도를 낮추고 감사 가능성을 제한합니다. MLLM은 강력한 추론 능력을 보여주지만, 종종 미세한 수준의 공간 이해와 객체 환각 문제에 어려움을 겪습니다. 기존 연구인 ByDeWay는 Layered-Depth-Based Prompting (LDP)이라는 훈련이 필요 없는 프레임워크를 소개하여 단안 심도 추정 기술을 사용하여 프롬프트를 구성함으로써 환각 현상을 완화합니다. 그러나 거친 심도 계층화는 동일한 기하학적 평면 내의 객체 간의 공간 관계, 예를 들어 투영(

Original Abstract

As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination. Prior work, ByDeWay, introduced Layered-Depth-Based Prompting (LDP), a training-free framework that mitigates hallucinations by structuring prompts using monocular depth estimation. However, coarse depth layering falls short in resolving object-to-object spatial relationships within the same geometric plane, such as projective ("left of", "above") and topological ("inside", "touching") relations. We propose ByDeWay-V2, which integrates explicit spatial relational context alongside depth cues, expressed as human-readable predicates that serve as auditable evidence for downstream decision support. Using an open-vocabulary object detector (YOLO-World-L), our framework computes pairwise geometric relations between detected objects and injects them as structured spatial predicates into the MLLM prompt, bridging 3D scene depth and 2D spatial semantics without any training. We evaluate ByDeWay-V2 on the Visual Spatial Reasoning (VSR) and BLINK benchmarks across multiple MLLMs, with hallucination grounding assessed via POPE. On the BLINK spatial subset, ByDeWay-V2 achieves a 46 percent relative F1 improvement over LDP for Qwen2.5-VL, and recovers BLIP-Base's spatial reasoning on VSR from near-random performance to a competitive F1 of 0.53. Our lightest configuration operates under a strict 40-token context budget on CPU, showing the framework's suitability for resource-constrained, real-time decision-support settings.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!