2608.01437v1 Aug 02, 2026 cs.AI

라우팅 포화 너머: 다중 모드 연속적인 지시 조정(Multimodal Continual Instruction Tuning)에서의 전문가 라우팅에 대한 장기적이고 점진적인 관점

Beyond Routing Saturation: A Long-Horizon Class-Incremental Perspective on Expert Routing in Multimodal Continual Instruction Tuning

Furao Shen
Furao Shen
Citations: 65
h-index: 5
Huiyu Yi
Huiyu Yi
Citations: 5
h-index: 2
Zhiming Xu
Zhiming Xu
Citations: 10
h-index: 2
Dunwei Tu
Dunwei Tu
Citations: 2
h-index: 1
Baile Xu
Baile Xu
Citations: 223
h-index: 6
Bogang Zhang
Bogang Zhang
Citations: 0
h-index: 0
Zhen-Hao Xie
Zhen-Hao Xie
Citations: 0
h-index: 0
Yongqi Xu
Yongqi Xu
Citations: 0
h-index: 0

다중 모드 연속적인 지시 조정(MCIT)은 다중 모드 대규모 언어 모델이 새로운 작업을 순차적으로 학습하면서 이전에 습득한 기능을 유지할 수 있도록 합니다. 많은 최근 연구에서는 작업별 LoRA 전문가를 사용하고 추론 시 각 입력을 하나 이상의 전문가에게 라우팅합니다. 그러나 전문가 라우팅의 근본적인 과제인 작업 식별 문제는 아직 충분히 탐구되지 않았습니다. 우리는 널리 사용되는 MCIT 벤치마크에서 라우팅이 거의 포화 상태에 있음을 보여줍니다. 텍스트 기반의 특징(fingerprint)이 작업 ID를 유출하고, 경쟁하는 전문가가 몇 개밖에 없는 짧은(4~10개 작업) 시퀀스는 장기적인 라우팅 문제를 가립니다. 이러한 과제를 드러내기 위해, 우리는 약화된 텍스트 기반 특징을 가진 34개의 작업으로 구성된 장기적인 MCIT 벤치마크인 FLEX(Fingerprint-reduced Long-horizon Expert eXamination)를 소개합니다. FLEX는 유사한 지시 및 답변 형식을 갖지만 다양한 시각적 및 지식 도메인을 가진 작업을 그룹화하고, 외부 템플릿을 정규화하며, 훨씬 더 큰 전문가 풀에서 라우팅 성능을 평가합니다. 더욱 중요한 점은, 우리는 점진적인 LoRA 라우팅을 소프트한 작업-클래스 다중 모드 점진적 학습(MCIL) 문제로 공식화했습니다. 각 작업은 증분적인 라우팅 클래스를 정의하며, 이 클래스의 전체 스코어 분포는 LoRA 혼합 가중치를 제공하고, 하드 라우팅은 이러한 이산적인 특수한 경우입니다. FLEX는 확장되는 작업 식별 문제를 드러내고 있으며, MCIL 공식화는 기존의 점진적 학습 방법을 전문가 라우팅에 적용하기 위한 체계적인 인터페이스를 제공합니다. 우리는 PureLoRA를 제어된 기준선으로 사용하고, 네 가지 CIL 방법을 네 가지 MCIT 프레임워크에 적용하여 LoRA 전문가나 생성 파이프라인을 수정하지 않고 성능을 향상시켰습니다. 우리의 플러그인 라우터는 엄격한 LoRA 일치율을 최대 16.3%p만큼, 전체 MacroScore를 최대 4.6 포인트만큼 향상시킵니다. 코드는 다음 위치에서 확인할 수 있습니다: https://github.com/RINC-CL/FLEX

Original Abstract

Multimodal Continual Instruction Tuning (MCIT) enables multimodal large language models to acquire new tasks sequentially while retaining previously learned capabilities. Many recent methods maintain task-specific LoRA experts and route each input to one or more experts at inference. Yet the task-identification problem underlying expert routing remains under-explored. We show that routing is nearly saturated on widely used MCIT benchmarks. Textual fingerprints that leak task identity and short 4--10-task sequences with few competing experts jointly obscure the long-horizon routing problem. To expose this challenge, we introduce FLEX (Fingerprint-reduced Long-horizon Expert eXamination), a 34-task long-horizon MCIT benchmark with weakened textual fingerprints. FLEX groups tasks with similar instruction and answer formats but diverse visual and knowledge domains, normalizes their outer templates, and evaluates routing over a substantially larger expert pool. Crucially, we formulate progressive-LoRA routing as soft task-as-class Multimodal Class-Incremental Learning (MCIL): each task defines an incremental routing class, whose complete score distribution supplies the LoRA mixture weights, with hard routing as a discrete special case. FLEX exposes this expanding task-identification challenge, while the MCIL formulation provides a principled interface for transferring CIL methods to expert routing. We instantiate PureLoRA as a controlled baseline and adapt four CIL methods to four MCIT frameworks without modifying their LoRA experts or generation pipelines. Our plug-in routers improve strict LoRA matching by up to 16.3 percentage points and overall MacroScore by up to 4.6 points. Code is available at: https://github.com/RINC-CL/FLEX

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!