LEDGERMIND: 구조화된 증거 장부를 활용한 출처 제약 하의 다중 모드 에이전트 추론
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
시각 질의 응답을 위한 다중 모드 에이전트는 점점 더 복잡한 단계적 과정을 거치며, 인지, 검색 및 추론이 상호 연결되지만, 평가 방법은 여전히 최종 답변 정확도에 크게 의존합니다. 이러한 전체적인 신호는 올바른 답변이 실제 근거를 통해 도출되었는지, 언어적 선입견으로 인해 얻어졌는지, 아니면 우연한 오류 제거 효과로 인해 발생했는지 판단하기 어렵습니다. 본 연구에서는 다중 모드 에이전트의 작동 과정을 출처 제약 하의 상태 머신으로 간주합니다. 각 단계의 결과는 구조화된 증거 장부로 정규화되어 전체 과정의 상태를 나타내며, 이후 추론 및 의사 결정은 해당 장부에 기록된 내용만을 참조할 수 있습니다. 또한, 개체 수준 및 숫자 수준에서 근거 검증을 수행하고, 오류 수정은 유형화된 상태 전이 과정을 통해 이루어지며, 이 과정에서 도구에서 생성된 출처 정보 없이 새로운 내용을 추가할 수 없습니다. 이러한 설계를 바탕으로 LedgerMind (구조화된 증거 장부를 활용한 출처 제약 하의 다중 모드 에이전트 추론)를 구현했으며, 여기에는 세 단계로 구성된 근거 프로토콜, 질문의 복잡성에 따라 추론 깊이를 조절하는 적응형 이중 경로 디스패처, 그리고 형식적인 출처 정보 증폭 방지 보장을 제공하는 이벤트 기반 검증 및 수정 엔진이 포함됩니다. LedgerMind를 사용하여 최종 답변 정확도로는 드러나기 쉬운 네 가지 일반적인 오류 패턴을 해결하고자 합니다. 이러한 패턴으로는 근거 없는 중간 추론, 참조를 통한 개체 환각 (Phantom Grounding), 단순 질의에 대한 과도한 추론, 그리고 수정 과정에서의 정보 증폭 등이 있습니다. 다양한 다중 모드 추론 벤치마크 및 기반 모델(MLLM)에 대한 실험 결과, LedgerMind는 답변 정확도와 전체 과정의 신뢰성을 모두 향상시키는 것으로 나타났습니다.
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.