2608.03660v1 Aug 04, 2026 cs.AI

암묵적인 요소를 제어하다: 지속적인 다중 모드 후속 학습을 위한 이중 채널 위험 인식 강화 학습 미세 조정

Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training

Yangyang Wu
Yangyang Wu
Citations: 609
h-index: 9
Mengying Zhu
Mengying Zhu
Citations: 7
h-index: 2
Tangyue Jin
Tangyue Jin
Citations: 0
h-index: 0
Meng Xi
Meng Xi
Citations: 20
h-index: 3
Yibei Liu
Yibei Liu
Citations: 0
h-index: 0
Jiajun Chen
Jiajun Chen
Citations: 0
h-index: 0
Qianle Zhang
Qianle Zhang
Citations: 0
h-index: 0

강화 학습 기반 미세 조정(RFT)은 다중 모드 대규모 언어 모델의 지속적인 후속 학습에서 고질적인 망각 현상에 대한 내성을 가지는 것으로 널리 알려져 있습니다. 그러나, 과도한 작업 분포 변화가 발생할 경우, 대표적인 RFT 알고리즘에서 발생하는 망각이 급격하게 증가합니다. 이는 RFT 본연에 내재된 암묵적인 보상-분산 정규화 때문에 발생하며, 통제되지 않는 최적화 위험을 억제하는 데 한계가 있습니다. 우리는 명시적인 위험 관리를 위한 최초의 이중 채널 프레임워크인 Risk-Aware Policy Optimization (RAPO)을 제안합니다. 정책 채널에서는, Risk-Aware Policy Scaling이 rollout 신뢰도와 Fisher 영감을 받은 지역 예측 감수성을 통해 샘플별 업데이트 크기를 적응적으로 조정합니다. 데이터 채널에서는, Risk-Aware Dynamic Bucket Sampling이 동적 위험 계층화를 통해 훈련 배치 재구성하여 최적화가 유익하면서도 안정적인 샘플을 향하도록 유도합니다. RAPO는 어떠한 작업 간의 메모리도 필요로 하지 않는 플러그 앤 플레이 방식으로 설계되었으며, 수정 없이 모든 RFT 알고리즘에 적용 가능합니다. 공개된 MLLM-CL 벤치마크에서, RAPO는 RLOO 기반 모델보다 최종 망각을 79.8% 감소시키면서도 새로운 작업에서의 경쟁력을 유지했습니다.

Original Abstract

Reinforcement fine-tuning (RFT) is widely believed to inherently resist catastrophic forgetting in continual post-training of multimodal large language models. Under pronounced task distributional shifts, however, forgetting across representative RFT algorithms escalates sharply. This stems from the implicit reward-variance regularization inherent to RFT, which proves incapable of suppressing uncontrolled optimization risk. We propose Risk-Aware Policy Optimization (RAPO), the first dual-channel framework for explicit risk governance in continual RFT. On the policy channel, Risk-Aware Policy Scaling adaptively calibrates per-sample update magnitude via rollout reliability and Fisher-inspired local predictive sensitivity; on the data channel, Risk-Aware Dynamic Bucket Sampling reorganizes training batches through dynamic risk stratification, steering optimization toward informative yet stable samples. As a plug-and-play strategy requiring no cross-task memory, RAPO generalizes to any RFT algorithm without modification. On the public MLLM-CL benchmark, RAPO reduces final forgetting by 79.8% relative to its RLOO backbone while retaining new-task competitiveness.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!