2606.11853v1 Jun 10, 2026 cs.CV

작업 인지 구조화 메모리: 동적 다중 모달 컨텍스트 학습을 위한 방법

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning

Ziwei Chen
Ziwei Chen
Citations: 652
h-index: 5
Zhirui Chen
Zhirui Chen
Citations: 115
h-index: 4
Ling Shao
Ling Shao
Citations: 6
h-index: 1

다중 모달 대규모 언어 모델(MLLM)은 빠른 작업 적응을 위해 컨텍스트 학습(ICL)에 의존하지만, 유한한 컨텍스트 창과 긴 다중 모달 시퀀스에서 키-값(KV) 캐시의 증가하는 비용으로 인해 확장성이 심각하게 제한됩니다. 기존의 메모리 압축 방법은 일반적으로 경직된 토큰 제거 또는 샘플 종속 중요도 추정을 사용하는데, 이는 편향을 유발하고 의미 구조를 파괴하며, 특히 시각적 표현에 영향을 미치며, 새로운 쿼리에 적응할 수 없는 정적인 메모리를 생성합니다. 본 연구에서는 이러한 한계를 극복하기 위해 작업 인지, 구조 보존, 동적 접근이 가능한 메모리 구성 방법을 제시하는 TASM(Task-Aware Structured Memory) 프레임워크를 소개합니다. TASM은 작업 벡터 기반의 압축을 사용하여 샘플별 신호를 작업 수준의 방향으로 대체함으로써 데모 간에 공유되는 관련성을 파악합니다. 또한, 기본 다양체를 보존하기 위해 이중 그래프 매칭을 통한 의미 기반 토큰 병합을 적용하여 파괴적인 가지치기 없이 토큰을 집계합니다. 마지막으로, TASM은 핵심 메모리와 잠재적 저장소로 구성된 계층 구조를 통해 메모리를 구성함으로써 쿼리 적응형 동적 검색을 용이하게 합니다. 실험 결과는 TASM이 높은 압축률에서도 우수한 성능을 유지하며 효율성과 적응성을 효과적으로 균형 있게 유지함을 확인합니다.

Original Abstract

Multi-modal large language models (MLLMs) depend on in-context learning (ICL) for rapid task adaptation, but their scalability is severely limited by finite context windows and the growing cost of key-value (KV) caches in long multi-modal sequences. Existing memory compression approaches typically rely on rigid token removal or sample-dependent importance estimation, which introduces bias, disrupts semantic structure, particularly for visual representations, and yields static memories that cannot adapt to new queries. We introduce TASM (Task-Aware Structured Memory), a training-free framework that addresses these limitations through task-aware, structure-preserving, and dynamically accessible memory construction. TASM employs task-vector guided compression to replace sample-specific signals with a task-level direction that captures shared relevance across demonstrations. To preserve the underlying manifold, it applies semantics-aware token merging via bipartite graph matching, aggregating tokens without destructive pruning. Finally, TASM structures memory into a hierarchy comprising a compact Core Memory and a Latent Bank, facilitating query-adaptive dynamic retrieval. Evaluations confirm TASM maintains high performance under heavy compression, effectively balancing efficiency with adaptability.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!