탐색, 지도 작성, 기억, 결정: 인체 기반 시각 언어 모델(VLM)이 안전에 중요한 상황에 적합한가?
Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?
본 연구는 부분적인 정보만 주어졌을 때, 호기심으로 작동하는 시각-언어 모델(VLM)의 공간 이해 능력을 '공간 이론 프레임워크'(ToS)를 통해 평가합니다. 인공지능 기술이 안전에 중요한 상황에 점점 더 많이 적용됨에 따라, VLM이 견고한 공간 기억 능력을 가지고 있으며 신뢰할 수 있는 결정을 내릴 수 있는지 여부를 파악하는 것이 중요합니다. 본 논문에서는 VLM의 결정이 물리적 증거에 기반하는지, 아니면 시각-언어 편향으로 인해 왜곡되는지를 평가하고, 그들의 기억 과정이 인간 인지 패턴과 일치하는지, 그리고 환경 위험 요소에 어떻게 반응하는지를 분석합니다. 우리는 ToS 프레임워크를 안전에 중요한 목표 지향적인 파이프라인인 '탐색, 지도 작성, 기억, 결정'(EMRD)으로 확장했습니다. 탐색 능력(Explore)은 환경 범위와 시간 효율성 지표를 통해 정량화하고, 공간 정확도(Map)는 다양한 심리적 지표를 사용하여 평가하며, 기억 지속성(Remember)은 일련의 심리적 지표를 통해 평가하고, 인지적 의사 결정(Decide)은 특정 지점 중심 지표를 사용하여 측정합니다. 연구 결과는 VLM이 의사 결정 능력 측면에서 훈련된 텍스트 정보에 기반하여 대피 지점을 자주 선택하지만, 그러한 선택을 정당화할 수 있는 공간적 이해력이 부족하다는 것을 보여줍니다. 또한, 낮은 조명 조건에서는 공간 추론 능력이 저하되지만, 질감 및 색상 변경에는 영향을 받지 않는다는 것을 확인했습니다. 이러한 결과는 VLM의 기억이 근본적으로 인간 인지에 비해 다르며, 예측 불가능한 위험을 초래할 수 있음을 시사합니다.
Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI techniques are increasingly applied to safety-critical scenarios, it is crucial to understand whether VLMs possess robust spatial memory and make reliable decisions. In this paper, we assess whether VLMs' decisions are based on physical evidence or are corrupted by visual-language biases, if their memory processes align with human cognitive patterns, and how they respond to environmental hazards. We extend the ToS framework into a safety-critical, goal-driven pipeline, named Explore, Map, Remember, and Decide (EMRD). We then quantify Exploration Competence (Explore) through metrics of environmental coverage and temporal efficiency, assess Spatial Fidelity (Map), evaluate, with a suite of psychological metrics, Memory Persistence (Remember), and measure, using focal-point metrics, Cognitive Decision-Making (Decide). Our results show that in terms of decision-making capabilities, VLMs frequently select evacuation points based on pre-trained textual priors while lacking the spatial grounding to justify their choices. We also show that spatial reasoning degrades in low-light conditions, but it is not affected by texture and colour tampering. Our findings suggest that VLM memory fundamentally diverges from human cognition, creating unpredictable risks of misalignment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.