2601.15549v1 Jan 22, 2026 cs.CV

VIOLA: 최소한의 주석을 활용한 비디오 기반 문맥 학습을 향하여

VIOLA: Towards Video In-Context Learning with Minimal Annotations

Ryo Fujii
Ryo Fujii
Keio University
Citations: 208
h-index: 9
Hideo Saito
Hideo Saito
Citations: 94
h-index: 5
Ryo Hachiuma
Ryo Hachiuma
Citations: 504
h-index: 11

멀티모달 대규모 언어 모델(MLLM)을 새로운 비디오 도메인으로 일반화하는 것은 실제 적용에 필수적이지만, 레이블이 지정된 데이터의 부족으로 인해 여전히 어려운 과제입니다. 문맥 학습(ICL)은 학습 없이 적응할 수 있는 방법을 제공하지만, 일반적인 방법은 대규모의 주석이 필요하며, 이는 전문가의 주석이 필요한 산업 또는 수술 환경과 같은 특수 환경에서는 비현실적입니다. 이러한 격차를 해소하기 위해, 우리는 최소한의 주석을 활용한 비디오 기반 문맥 학습 프레임워크인 VIOLA(Video In-cOntext Learning with minimal Annotation)를 소개합니다. 첫째, 제한된 주석 예산을 최대한 활용하기 위해, 밀도-불확실성 가중 샘플링을 제안합니다. 표준적인 다양성 또는 불확실성 전략은 시각적 이상치를 선택할 위험이 있지만, 우리의 방법은 밀도 추정을 활용하여 동시에 다양하고, 대표적이며, 정보적인 샘플을 식별합니다. 둘째, 노이즈 전파를 방지하면서 나머지 레이블이 없는 데이터를 활용하기 위해, 하이브리드 풀을 구성하고, 신뢰도 기반 검색 및 신뢰도 기반 프롬프팅을 도입합니다. 이러한 메커니즘은 레이블의 신뢰성을 명시적으로 모델링하여, 유사성 및 신뢰도의 복합 점수를 기반으로 데모를 검색하고, MLLM이 검증된 정답과 노이즈가 포함된 가짜 레이블을 적응적으로 구별할 수 있도록 합니다. 네 가지 MLLM을 사용하여 아홉 가지 다양한 벤치마크에서 수행한 광범위한 실험 결과, 우리의 프레임워크는 저자원 환경에서 다양한 기본 모델보다 훨씬 뛰어난 성능을 보이며, 최소한의 주석 비용으로 강력한 적응성을 달성합니다.

Original Abstract

Generalizing Multimodal Large Language Models (MLLMs) to novel video domains is essential for real-world deployment but remains challenging due to the scarcity of labeled data. While In-Context Learning (ICL) offers a training-free adaptation path, standard methods rely on large annotated pools, which are often impractical in specialized environments like industrial or surgical settings since they require the experts' annotations. To bridge this gap, we introduce VIOLA (Video In-cOntext Learning with minimal Annotation), a label-efficient framework that synergizes minimal expert supervision with abundant unlabeled data. First, to maximize the efficiency of a strict annotation budget, we propose density-uncertainty-weighted sampling. Unlike standard diversity or uncertainty strategies that risk selecting visual outliers, our method leverages density estimation to identify samples that are simultaneously diverse, representative, and informative. Second, to utilize the remaining unlabeled data without noise propagation, we construct a hybrid pool and introduce confidence-aware retrieval and confidence-aware prompting. These mechanisms explicitly model label reliability, retrieving demonstrations based on a composite score of similarity and confidence while enabling the MLLM to adaptively distinguish between verified ground truths and noisy pseudo-labels. Extensive experiments across nine diverse benchmarks using four MLLMs demonstrate that our framework significantly outperforms various baselines in low-resource settings, achieving robust adaptation with minimal annotation costs.

2 Citations
0 Influential
5.5 Altmetric
29.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!