기억이 거짓말을 할 때: VLM 에이전트의 공간 기억 노후화에 대한 실증적 연구
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
메모리 기반 VLM(Vision-Language Model) 에이전트는 지속적인 공간 지식을 활용하지만, 환경 변화로 인해 해당 지식은 눈치채지 못한 채 점차 노후화됩니다. 본 연구에서는 에이전트가 확신에 찬 기억 정보와 상반되는 관찰 결과를 어떻게 처리하는지, 그리고 현재 모델이 이러한 충돌을 안전상의 문제로 이어지기 전에 감지할 수 있는지 묻습니다. 동적인 FrozenLake 테스트 환경을 사용하여, 세 가지 폐쇄형 모델과 세 가지 공개 가중치 VLM에 대해 텍스트와 이미지 입력을 모두 사용(총 1,800개의 노후화 감지 실행 및 4가지 LLM 내비게이터를 사용하는 12,000개의 텍스트 모드 내비게이션 에피소드, 각 모델당 50개의 시드 설정)하여 노후화 감지 작업과 다운스트림 내비게이션 작업을 결합했습니다. 세 가지 주요 결과가 도출되었습니다. 첫째, 텍스트 기반 해결 가능성이 반드시 시각적 이해를 의미하지는 않습니다. 텍스트를 통해 노후화된 항목을 안정적으로 식별하는 모델이라 할지라도 동일한 그리드에서 시각적 F1 점수가 0.887에서 0.067까지 다양하며, 성능이 가장 낮은 모델은 지속적으로 이미지 정보를 무시하는 자연스럽고 확신에 찬 결정을 내립니다. 둘째, 검토 없이 노후화된 메모리를 사용하는 것은 안전상의 위험을 초래합니다. 주요 GPT-4o 설정에서, 원시 메모리에 의존하는 에이전트는 아무런 메모리도 제공받지 않은 동일한 에이전트보다 훨씬 더 자주 실패합니다. 셋째, 검토는 도움이 되지만 격차를 완전히 해소하지 못합니다. 투명한 실행 시간 필터는 텍스트 모드에서 상당 부분의 안전 위험을 줄이지만, 심지어 완벽하게 정확한 노후화 정보 라벨도 현재 그리드 크기에서는 더 큰 향상을 가져다주지 못하며, 시각적 검토가 신뢰할 수 없을 때는 필터링이 일관된 이점을 제공하지 않습니다. 종합적으로 이러한 결과는 공간 기억의 노후화를 안전상의 문제로 규정하고, 메모리 기반 에이전트에서 메모리와 관찰 간의 충돌 상황에서 안정적인 시각적 이해와 행동 선택을 핵심 과제로 제시합니다.
Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens when an agent must reconcile a confident memory claim with a contradicting observation, and whether current models can catch the conflict before it becomes a safety-relevant mistake. Using a dynamic FrozenLake testbed, we pair a staleness-detection task with a downstream navigation task across three closed-source models and three open-weight VLMs under both text and image inputs (1,800 detection runs, and 12,000 text-mode navigation episodes over four LLM navigators at a shared 50-seed scale). Three findings emerge. First, text solvability does not imply visual grounding: models that flag stale entries reliably from text nonetheless span vision F1 from 0.887 down to 0.067 on the identical grids, and the weakest keeps making fluent, confident decisions that ignore the image. Second, consuming stale memory without an audit is a safety liability: in our primary GPT-4o setting, an agent that trusts raw memory dies more than twice as often as the same agent given no memory at all. Third, auditing helps but does not close the gap: a transparent read-time filter removes much of the safety cost in text mode, yet even oracle stale labels bring no further significant gain on the current grid size, and when visual auditing is unreliable, filtering yields no consistent benefit. Together these results frame spatial-memory staleness as a safety failure mode and isolate reliable visual grounding and action selection under memory--observation conflict as the central open challenges for memory-augmented agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.