2607.28678v1 Jul 29, 2026 cs.AI

ViSAGE: 장편 비디오 이해를 위한 자기 교정 메모리 구축

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

Guanjie Cheng
Guanjie Cheng
Citations: 34
h-index: 3
Xinkui Zhao
Xinkui Zhao
Citations: 32
h-index: 3
Yifan Zhang
Yifan Zhang
Citations: 21
h-index: 3
Enbo Chen
Enbo Chen
Citations: 0
h-index: 0
Yueshen Xu
Yueshen Xu
Citations: 29
h-index: 3
Chang Liu
Chang Liu
Citations: 305
h-index: 7
Naibo Wang
Naibo Wang
Citations: 76
h-index: 5

장기적인 환경에서 작동하는 다중 모드 에이전트는 개체 일관성을 유지하고 시간적 맥락을 고려한 추론을 지원하기 위해 멀티미디어 메모리를 구축하고 지속적으로 업데이트해야 합니다. 그러나 기존의 에이전트 기반 메모리 접근 방식은 종종 공격적인 압축 및 세그먼트별 처리 과정에서 미세한 개체 정보를 손실합니다. 또한, 이러한 방식은 벡터 유사성 검색에 크게 의존하는데, 이는 의미적으로 관련 있지만 개체가 일치하지 않는 정보를 가져올 수 있으며, 이로 인해 개체 혼동, 오류 전파 및 환각적인 답변이 발생할 수 있습니다. 본 논문에서는 자기 교정 기능을 갖춘 개체 중심 메모리를 구축하는 다중 모드 에이전트 기반 메모리 프레임워크인 ViSAGE를 제안합니다. 특히, ViSAGE는 장기간에 걸쳐 크로스-모달 바인딩을 통해 개체의 동일성을 고정합니다. 그런 다음 양방향 메모리 정제 과정을 적용하여 지연된 개체 정보를 전파하고, 과거 기록을 사후적으로 통합하며, 향후 추론 능력을 향상시킵니다. 또한, 다중 에이전트 간의 교차 검증을 도입하여 검색된 증거가 개체-증거 일관성 제약 조건을 충족하는지 평가합니다. 이를 통해 증거가 부족할 경우 답변하지 않고, 근거 없는 답변을 제공하는 것을 방지할 수 있습니다. 광범위한 실험 결과에서 ViSAGE는 가장 강력한 기준 모델보다 5.9% 더 높은 정확도를 달성하며 일관되게 우수한 성능을 보임을 입증했습니다.

Original Abstract

Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive compression and segment-wise processing. They also rely heavily on vector similarity retrieval, which can surface semantically related yet identity-mismatched evidence, leading to entity confusion, error propagation, and hallucinated answers. We propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories. Specifically, ViSAGE anchors entity identity via cross-modal binding over long temporal ranges. It then applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning. We also introduce multi-agent cross-verification to assess retrieved evidence under an identity-evidence alignment onstraint, enabling abstention instead of unsupported answers when evidence is missing. Extensive results demonstrate that ViSAGE consistently outperforms the strongest baseline, achieving 5.9% higher accuracy.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!