2607.08497v1 Jul 09, 2026 cs.CV

인지 구조 기반의 다중 모드 에이전트: 다중 모드 이해, 생성 및 편집을 위한 시스템

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

Ge Li
Ge Li
Citations: 230
h-index: 5
Chen Li
Chen Li
Citations: 67
h-index: 6
Jing Lyu
Jing Lyu
Citations: 13
h-index: 2
Feng Wang
Feng Wang
Citations: 6
h-index: 1
Canmiao Fu
Canmiao Fu
Citations: 48
h-index: 5
Zhipeng Huang
Zhipeng Huang
Citations: 48
h-index: 5

최근 통합 다중 모드 모델은 단일 아키텍처로 시각/언어 이해와 이미지 생성/편집을 동시에 수행할 수 있음을 보여줍니다. 그러나 이러한 모델들은 공유 컨텍스트 윈도우에 모든 과거 시각적 및 텍스트 입력을 반복적으로 입력하는데, 이는 시각 토큰 폭증과 신뢰성이 떨어지는 대화 턴 간 참조로 인해 장기적인 다중 모드 대화 능력을 제한합니다. 본 연구에서는 시각 정보를 에피소딕 시각 기억(Episodic Visual Memory)에 외부 저장하고 추론 과정에서 관련된 에피소드를 선택적으로 활성화하는 인지 구조 기반의 다중 모드 에이전트를 제안합니다. 이 에이전트는 구조화된 시각적 추상화를 수행하는 지각 추상 엔진(Perceptual Abstraction Engine), 턴 간 기억 검색을 위한 인지 검색 엔진(Cognitive Retrieval Engine), 그리고 자율적인 작업 추론 및 행동 계획을 수행하는 다중 모드 실행 제어기(Multimodal Executive Controller)로 구성됩니다. 기존 데이터셋에 존재하는 턴 단위의 검색 지도 학습 부족 문제를 해결하기 위해, 우리는 구조화된 멀티 턴 대화를 프로그래밍 방식으로 생성하고 세분화된 검색 주석을 제공하는 통합 시나리오 엔진(Unified Scenario Engine)을 개발했습니다. 이를 통해 강화 학습을 활용하여 추상화 및 검색 정책을 최적화할 수 있습니다. 또한, 에피소딕 시각 기억 능력을 평가하기 위해 난이도별로 구성된 장기적인 시각-대화 벤치마크를 구축했습니다. 제안하는 8B 에이전트는 20턴 대화에서 91.4%의 검색 정확도를 달성하며, 32B 기준 모델보다 +8.2% 향상된 성능을 보이고, 턴당 추론 시간을 거의 절반(23.1초 -> 12.7초)로 단축했습니다. 또한, 지속적인 다중 모드 기억, 웹 접근, 이미지 생성/편집/조합 도구, 그리고 OpenAI 호환 서버를 통합하는 인지 구조 기반의 다중 모드 에이전트 플랫폼(Cognitive-structured Multimodal Agent Harness, CMA-Harness)을 개발하여 실제 활용 가능성을 제시합니다. 구조화된 기억과 모듈식 의사 결정은 단일 파라미터 확장 방식보다 장기적인 다중 모드 에이전트에 대한 더욱 확장 가능하고 효율적인 패러다임을 제공합니다. 코드: https://github.com/caseclose/cma-harness ; 프로젝트 페이지: https://caseclose.github.io/cma-harness/

Original Abstract

Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a shared context window, limiting long-horizon multimodal dialogue due to visual token explosion and unreliable cross-turn referencing. We propose a Cognitive-structured Multimodal Agent that externalizes visual information into an Episodic Visual Memory and selectively reactivates relevant episodes during reasoning. The agent consists of a Perceptual Abstraction Engine for structured visual abstraction, a Cognitive Retrieval Engine for cross-turn memory retrieval, and a Multimodal Executive Controller for autonomous task inference and action planning. To address the lack of turn-level retrieval supervision in existing datasets, we develop a Unified Scenario Engine that programmatically generates structured multi-turn conversations with fine-grained retrieval annotations, enabling reinforcement learning to optimize abstraction and retrieval policies. We also construct a long-horizon visual-dialogue benchmark stratified by difficulty to evaluate episodic visual recall. Our 8B agent achieves 91.4% retrieval accuracy over 20-turn sessions, surpassing 32B baselines by +8.2% while nearly halving per-turn inference time (23.1s -> 12.7s). We further present the Cognitive-structured Multimodal Agent Harness (CMA-Harness), a tool-augmented deployment of the same cognitive structure integrating persistent multimodal memory, web access, image generation/editing/composition tools, and OpenAI-compatible serving. Structured memory and modular decision-making offer a more scalable, efficient paradigm for long-horizon multimodal agents than monolithic parameter scaling. Code: https://github.com/caseclose/cma-harness ; Project page: https://caseclose.github.io/cma-harness/

0 Citations
0 Influential
33.986122886681 Altmetric
0.0 Score
Original PDF
8

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!