MathVis-Fine: 점진적인 의존성 기반 학습을 통한 시각적 감독과 필요성의 정렬 - 다중 모드 수학 문제 해결
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning
체인 오브 소트 (Chain-of-Thought, CoT) 추론은 순수 언어 영역에서 다중 모드 시나리오로 확장되었지만, 기존 접근 방식은 종종 시각적 입력을 동질적이거나 보조적인 신호로 취급하여 수학 문제 해결 과정에서의 텍스트와 이미지 간의 복잡하고 샘플별 의존성을 제대로 파악하지 못합니다. 이러한 문제는 두 가지 핵심 문제점을 야기합니다: 첫째, 시각 콘텐츠에 대한 감독 신호가 일반화되고 거칠어, 각 샘플에서 필요한 시각 정보의 실제 중요성에 대한 적응력이 부족합니다. 둘째, 시각적 보상이 입력 간의 상호 보완적인 관계를 구별하지 않고 균일하게 적용될 때, 학습 피드백이 부정확해집니다. 이러한 한계는 모델이 정확한 다중 모드 추론을 달성하는 것을 방해합니다. 본 연구에서는 수학적 추론에서 미세한 시각적 의존성을 모델링하기 위한 프레임워크를 제안합니다. 먼저, 미세한 시각적 주석과 함께 시각적 의존성 등급이 포함된 MathVis-Fine 데이터셋을 구축했습니다. 이 데이터셋을 기반으로, 각 샘플의 고유한 시각적 의존성 수준에 따라 정답 정확도 보상과 시각적 연결 보상을 균형 있게 조정하는 두 단계의 점진적인 시각적 강화 학습 패러다임을 도입하여 보상 편향을 완화하고 감독 정확도를 향상시킵니다. 광범위한 실험 결과, MathVis-Fine 프레임워크가 시각적 의존성을 기반으로 시각적 인지 능력을 효과적으로 점진적으로 향상시키며, 다중 모드 수학 문제 해결을 위한 더욱 정교한 학습 프레임워크를 제공한다는 것을 보여줍니다. 본 연구의 데이터셋은 논문 게재 후 공개될 예정입니다.
Chain-of-Thought (CoT) reasoning has extended from purely linguistic domains to multimodal scenarios; however, existing approaches often treat visual inputs as homogeneous or auxiliary signals, failing to capture the intricate and sample-specific dependencies between text and images in mathematical problem-solving. This gives rise to two core issues: first, the supervisory signals for visual content are generalized and coarse-grained, lacking adaptation to the actual necessity of visual information in each sample; second, training feedback becomes inaccurate when visual rewards are uniformly applied without distinguishing the complementary relationships among inputs. These limitations hinder models from achieving precise multimodal reasoning. In this work, we propose a framework for modeling fine-grained visual dependencies in mathematical reasoning. We first construct the MathVis-Fine dataset, augmenting fine-grained visual annotations with visual dependency ratings. Building upon this dataset, we introduce a two-stage progressive visual enhancement training paradigm that balances answer correctness rewards and visual grounding rewards according to the intrinsic visual dependency level of each sample, thereby mitigating reward bias and improving supervision accuracy. Extensive experiments demonstrate that the MathVis-Fine framework effectively enhances visual perception progressively based on visual dependency, offering a more precise training framework for multimodal mathematical reasoning. We will release the dataset upon acceptance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.