CURV: 커리큘럼 기반 시각적 근거 추론을 통한 차트 이해도 향상
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
차트 질의 응답(CQA)은 다중 모드 대규모 언어 모델(MLLM)이 시각적 이해와 논리적 추론을 통합하도록 요구하지만, 현재 모델들은 정확한 시각적 근거 설정 및 일관성 있는 추론 체인 구축에 어려움을 겪고 있습니다. 외부적인 체인 오브 소트 프롬프트 및 시각적 단서는 성능 향상에 크게 기여하지만, 현재 MLLM은 내재적인 시각적 근거 추론 능력이 부족하여 부정확한 인지와 시각적 증거와 동떨어진 추론을 야기합니다. 이러한 한계를 극복하기 위해, 우리는 CURV라는 커리큘럼 학습 프레임워크를 제안합니다. CURV는 CQA 문제를 다단계 시각적 근거 추론으로 재구성하여 내재적인 시각적 추론 능력을 개발하며, 각 단계에서 공간적 주의 집중을 통해 논리적 추론과 동적인 시각적 근거 설정을 조율합니다. 모델 학습을 돕기 위해, 우리는 다양한 차트 유형 및 추론 패턴에 대한 확장 가능한 합성 데이터 생성 기능을 갖춘 세 단계 커리큘럼 데이터셋인 CCQA를 추가로 소개합니다. 우리의 커리큘럼은 기본 단일 작업 추론부터 복잡한 다중 차트 조합 작업까지 체계적으로 진행됩니다. 실험 결과, CURV는 기존 모델 대비 최대 20.50%의 성능 향상을 달성했으며, 실제 환경 벤치마크(최대 12.30% 향상) 및 외부 도메인의 다중 모드 추론 작업(최대 10.20% 향상)에도 적용 가능하여, 동적인 근거 설정을 통한 시각적 추론 능력을 내재화함으로써 차트 이해도를 향상시키는 데 효과적임을 입증했습니다. 코드: https://xhguo7.github.io/CURV/
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain-of-thought prompting and visual cues significantly improve performance, current MLLMs lack intrinsic visual grounded reasoning capabilities, leading to inaccurate perception and reasoning disconnected from visual evidence. To address these limitations, we propose CURV, a curriculum learning framework that develops intrinsic visual reasoning capabilities by reformulating CQA as multi-step visual grounded reasoning, where each step coordinates logical reasoning with dynamic visual grounding through spatial attention concentration. To assist model learning, we further introduce CCQA, a three-level curriculum dataset with scalable synthetic generation across diverse chart types and reasoning patterns. Our curriculum systematically progresses from basic single-operation reasoning to complex multi-chart compositional tasks. Experiments demonstrate that CURV achieves up to $\uparrow20.50\%$ improvements over baselines and is generalizable to real-world benchmarks (up to $\uparrow12.30\%$) and out-of-domain multimodal reasoning tasks (up to $\uparrow10.20\%$), validating the effectiveness of internalizing visual reasoning with dynamic grounding for enhanced chart understanding capabilities. Code is available at: https://xhguo7.github.io/CURV/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.