2608.09227v1 Aug 10, 2026 cs.AI

Omni2LoRA: 일관성 유지 파라미터 메모리 - 효율적인 올인원 언어 모델을 위한 기술

Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models

Manan Suri
Manan Suri
Citations: 166
h-index: 7
Dinesh Manocha
Dinesh Manocha
Citations: 100
h-index: 6
Puneet Mathur
Puneet Mathur
University of Maryland
Citations: 1,180
h-index: 17

옴니모달 언어 모델(OLM)은 통합된 오디오-비주얼 이해를 가능하게 하지만, 긴 결합 토큰 시퀀스를 처리하는 것은 계산 비용 측면에서 매우 비효율적입니다. 최근의 토큰 압축 방법들은 이러한 부담을 줄이려고 시도하지만, 모달리티를 개별적으로 압축하면 일관성 있는 추론에 필수적인 시간적 크로스-모달 연결 고리가 손실되는 경우가 많습니다. 본 논문에서는 일관성을 유지하는 컨텍스트 증류를 통해 효율적인 파라미터 메모리 압축을 위한 2단계 프레임워크인 Omni2LoRA를 소개합니다. 이 방법은 토큰 병목 현상을 완전히 우회합니다. 먼저, Perceiver 하이퍼 네트워크는 동결된 OLM에서 얻은 중간 표현을 처리하여 멀티모달 컨텍스트를 단일 순방향 패스에서 풀랭크 저차 적응(LoRA) 어댑터로 인코딩합니다. 결과적으로 발생하는 파라미터 크기가 녹음 길이에 따라 선형적으로 증가하는 것을 방지하기 위해, 본 논문에서는 그룹 상대 정책 최적화(GRPO)를 통해 이산적인 랭크 할당 정책을 최적화합니다. GRPO는 모달리티를 제거한 가상 보상을 사용하여 오디오-비주얼 일관성의 손실을 명시적으로 처벌하여 모델이 고정된 하위 선형 랭크 예산을 분리된 시각적 특징 대신 시너지 효과가 있는 크로스-모달 연결 지점에 할당하도록 합니다. 세 가지 옴니모달 백본에서, Omni2LoRA는 전체 컨텍스트 추론과 강력한 토큰 압축 기준(OmniZip, OMAC, O-MARC)보다 30%의 랭크 예산으로 작동하면서 네 가지 오디오-비주얼 질의응답 벤치마크에서 평균 정확도를 8~12% 향상시켰습니다. 또한 토큰 잘라내기 방법이 급격하게 성능 저하를 보이는 75%에 이르는 압축 비율에서도 안정적인 성능을 유지했습니다. 본 연구는 멀티모달 메모리를 고정된 예산을 가진 재사용 가능한 파라미터 상태로 변환하여, 질의 응답 시 멀티모달 토큰 로드를 0으로 줄이고 전체 컨텍스트 추론에 비해 최대 12배 빠른 시간(TTFT)을 제공하며, 몇 번의 질의 후에는 0.5초 미만의 처리 시간을 달성합니다.

Original Abstract

Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent token compression methods attempt to alleviate this burden, compressing modalities in isolation often destroys the temporal cross-modal anchors necessary for coherent reasoning. We introduce Omni2LoRA, a two-stage framework for efficient parametric memory compression via coherence-preserving context distillation that bypasses the token bottleneck entirely. First, a Perceiver hypernetwork processes intermediate representations from a frozen OLM to encode the multimodal context into a full-rank Low-Rank Adaptation (LoRA) adapter in a single forward pass. To prevent the resulting parameter footprint from scaling linearly with recording length, we optimize a discrete rank allocation policy via Group Relative Policy Optimization (GRPO) that uses a modality-ablated counterfactual reward to explicitly penalize the loss of audio-visual coherence, forcing the model to allocate its fixed sub-linear rank budget to synergistic cross-modal anchors rather than isolated visual features. Across three omnimodal backbones, Omni2LoRA operating at a 30% rank budget outperforms direct full-context inference and strong token-compression baselines (OmniZip, OMAC, O-MARC) on four audio-visual question answering benchmarks, improving average accuracy by 8-12% over the strongest baseline and remaining stable under compression ratios as tight as 75%, where token-pruning methods degrade sharply. By converting multimodal memory into a fixed-budget, reusable parameter state, our method drives answer-time multimodal-token load to zero, cutting per-query Time to First Token (TTFT) by up to 12x relative to full-context inference and amortizing to under 0.5s after a handful of queries.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!