암묵적인 요소를 제어하다: 지속적인 다중 모드 후속 학습을 위한 이중 채널 위험 인식 강화 학습 미세 조정
Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training
강화 학습 기반 미세 조정(RFT)은 다중 모드 대규모 언어 모델의 지속적인 후속 학습에서 고질적인 망각 현상에 대한 내성을 가지는 것으로 널리 알려져 있습니다. 그러나, 과도한 작업 분포 변화가 발생할 경우, 대표적인 RFT 알고리즘에서 발생하는 망각이 급격하게 증가합니다. 이는 RFT 본연에 내재된 암묵적인 보상-분산 정규화 때문에 발생하며, 통제되지 않는 최적화 위험을 억제하는 데 한계가 있습니다. 우리는 명시적인 위험 관리를 위한 최초의 이중 채널 프레임워크인 Risk-Aware Policy Optimization (RAPO)을 제안합니다. 정책 채널에서는, Risk-Aware Policy Scaling이 rollout 신뢰도와 Fisher 영감을 받은 지역 예측 감수성을 통해 샘플별 업데이트 크기를 적응적으로 조정합니다. 데이터 채널에서는, Risk-Aware Dynamic Bucket Sampling이 동적 위험 계층화를 통해 훈련 배치 재구성하여 최적화가 유익하면서도 안정적인 샘플을 향하도록 유도합니다. RAPO는 어떠한 작업 간의 메모리도 필요로 하지 않는 플러그 앤 플레이 방식으로 설계되었으며, 수정 없이 모든 RFT 알고리즘에 적용 가능합니다. 공개된 MLLM-CL 벤치마크에서, RAPO는 RLOO 기반 모델보다 최종 망각을 79.8% 감소시키면서도 새로운 작업에서의 경쟁력을 유지했습니다.
Reinforcement fine-tuning (RFT) is widely believed to inherently resist catastrophic forgetting in continual post-training of multimodal large language models. Under pronounced task distributional shifts, however, forgetting across representative RFT algorithms escalates sharply. This stems from the implicit reward-variance regularization inherent to RFT, which proves incapable of suppressing uncontrolled optimization risk. We propose Risk-Aware Policy Optimization (RAPO), the first dual-channel framework for explicit risk governance in continual RFT. On the policy channel, Risk-Aware Policy Scaling adaptively calibrates per-sample update magnitude via rollout reliability and Fisher-inspired local predictive sensitivity; on the data channel, Risk-Aware Dynamic Bucket Sampling reorganizes training batches through dynamic risk stratification, steering optimization toward informative yet stable samples. As a plug-and-play strategy requiring no cross-task memory, RAPO generalizes to any RFT algorithm without modification. On the public MLLM-CL benchmark, RAPO reduces final forgetting by 79.8% relative to its RLOO backbone while retaining new-task competitiveness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.