2602.11596v1 Feb 12, 2026 cs.AI

MAPLE: 모달리티 인식 사후 학습 및 학습 생태계

MAPLE: Modality-Aware Post-training and Learning Ecosystem

Nikhil Verma
Nikhil Verma
Citations: 88
h-index: 4
Manasa Bharadwaj
Manasa Bharadwaj
Citations: 17
h-index: 2
Kyung-Min Jin
Kyung-Min Jin
Citations: 66
h-index: 5
Kevin Ferreira
Kevin Ferreira
Citations: 55
h-index: 2
Konwoo Kim
Konwoo Kim
Citations: 45
h-index: 2
Youngjoon Kim
Youngjoon Kim
Citations: 16
h-index: 2
Minjun Kim
Minjun Kim
Citations: 7
h-index: 1
Jooyoung Yoo
Jooyoung Yoo
Citations: 8
h-index: 1

멀티모달 언어 모델은 이제 텍스트, 오디오, 비디오를 통합하여 통합 추론을 수행한다. 그러나 기존의 강화 학습(RL) 사후 학습 파이프라인은 모든 입력 신호를 동등하게 중요한 것으로 취급하며, 각 작업이 실제로 어떤 모달리티를 필요로 하는지는 무시한다. 이러한 모달리티를 고려하지 않는 학습은 정책 경사(policy-gradient) 분산을 증가시키고, 수렴 속도를 늦추며, 신호가 누락되거나 추가되거나 가중치가 재조정될 수 있는 실제 환경의 분포 변화에 대한 강건성을 저하시킨다. 우리는 다음과 같이 구성된 완전한 모달리티 인식 사후 학습 및 학습 생태계인 MAPLE을 소개한다. (1) 각 작업에 필요한 최소 신호 조합을 명시적으로 주석 처리한 최초의 벤치마크인 MAPLE-bench; (2) 이질적인 그룹 이점으로 인한 경사 분산을 줄이기 위해 모달리티 요구 사항별로 배치를 계층화하는 모달리티 인식 정책 최적화 프레임워크인 MAPO; (3) 더 어려운 신호 조합의 균형을 맞추고 우선순위를 부여하는 적응형 가중치 및 커리큘럼 스케줄링. 손실 집계, 클리핑, 샘플링 및 커리큘럼 설계에 대한 체계적인 분석을 통해 MAPO의 최적 학습 전략을 확립했다. 적응형 가중치와 커리큘럼 중심 학습은 신호 조합 전반의 성능을 더욱 향상시킨다. MAPLE은 단일/멀티모달 정확도 격차를 30.24% 줄이고, 3.18배 더 빠르게 수렴하며, 현실적으로 신호 접근이 제한된 상황에서도 모든 모달리티 조합에 걸쳐 안정성을 유지한다. MAPLE은 배포 준비가 된 멀티모달 RL 사후 학습을 위한 완전한 해법을 제시한다.

Original Abstract

Multimodal language models now integrate text, audio, and video for unified reasoning. Yet existing RL post-training pipelines treat all input signals as equally relevant, ignoring which modalities each task actually requires. This modality-blind training inflates policy-gradient variance, slows convergence, and degrades robustness to real-world distribution shifts where signals may be missing, added, or reweighted. We introduce MAPLE, a complete modality-aware post-training and learning ecosystem comprising: (1) MAPLE-bench, the first benchmark explicitly annotating minimal signal combinations required per task; (2) MAPO, a modality-aware policy optimization framework that stratifies batches by modality requirement to reduce gradient variance from heterogeneous group advantages; (3) Adaptive weighting and curriculum scheduling that balances and prioritizes harder signal combinations. Systematic analysis across loss aggregation, clipping, sampling, and curriculum design establishes MAPO's optimal training strategy. Adaptive weighting and curriculum focused learning further boost performance across signal combinations. MAPLE narrows uni/multi-modal accuracy gaps by 30.24%, converges 3.18x faster, and maintains stability across all modality combinations under realistic reduced signal access. MAPLE constitutes a complete recipe for deployment-ready multimodal RL post-training.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!