2607.25467v1 Jul 28, 2026 cs.CV

보여지고, 언급되고, 혹은 잊혀지는가? 다이얼로그 과정에서의 시각적 KV 메모리 효과 검증

Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns

Bo Wang
Bo Wang
Citations: 16
h-index: 2
Hong Chen
Hong Chen
Citations: 23
h-index: 3
Yuanlin Chu
Yuanlin Chu
Citations: 8
h-index: 2
Yuxuan Fan
Yuxuan Fan
Citations: 14
h-index: 2
Xuming Hu
Xuming Hu
Citations: 122
h-index: 6
Kangli Chen
Kangli Chen
Citations: 0
h-index: 0
Yubo Gao
Yubo Gao
Citations: 50
h-index: 4

상태를 유지하는 멀티모달 어시스턴트는 이미지를 한 번 인코딩하지만, 여러 턴 후에 그 이미지에 대한 질문에 답변할 수 있습니다. 어텐션 기반의 시각-KV 삭제(eviction)는 현재 관련 없는 정보는 앞으로도 불필요할 것이라고 가정합니다. 본 연구에서는 시각적 사실을 실제로 언제 안전하게 잊어버릴 수 있는지 질문하고, 시각적 메모리 효과 검증 프레임워크(Causal Visual Memory Audit, CVMA)를 소개합니다. CVMA는 쌍으로 연결된 단일-프레필 구조로, 특정 시각 영역, 전체 이미지 또는 이전 어시스턴트 텍스트가 사용할 수 없을 때 이후 답변이 어떻게 변하는지를 테스트합니다. VisDial 및 ConvBench 데이터셋에서, 현재 어텐션은 진단적인 주변 효용성(marginal utility) 제어를 통해 확인된 상당한 선택 여유 공간에도 불구하고 미래에 유용한 영역을 무작위 수준보다 못하게 평가하는 경우가 있습니다. 전체 점수는 이후 턴에서 비전 정보가 필요하지 않은 경우 이러한 실패를 가립니다. 하지만, 통제된 환경과 생성된 이력(history)은 어시스턴트 텍스트 KV가 이미지 KV를 대체하는 두 번째 방법을 보여줍니다. 이는 이미 언급되었지만 확실하게 확인되지 않은 사실에는 적용되지 않습니다. 테스트된 시스템에서 안전한 삭제는 낮은 미래 시각적 의존성 또는 사실에 특화된 언어 표현에 의해 지원됩니다. 즉, 현재 어텐션이 낮다고 해서 반드시 안전하게 삭제할 수 있는 것은 아닙니다.

Original Abstract

Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction assumes that evidence irrelevant now will remain dispensable, although future questions are unknown. We ask when a visual fact is actually safe to forget and introduce the Causal Visual Memory Audit (CVMA), a paired single-prefill framework that tests what later answers lose when a visual region, the whole image, or prior assistant text becomes unavailable. On VisDial and ConvBench, current attention can rank future-useful regions worse than random even though a diagnostic marginal-utility control shows substantial selection headroom. Aggregate scores hide this failure when later turns do not need vision; controlled and stock-generated histories reveal a second escape route, in which assistant-text KV replaces image KV for facts already stated but not reliably for unstated facts. In the tested stacks, safe forgetting is supported by low future visual dependence or fact-specific verbalization---not by low current attention.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!