2606.24477v1 Jun 23, 2026 cs.CV

video-SALMONN-R$^3$: 효율적인 비디오 이해를 위한 반복 시청, 재질문 및 재답변 학습

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Chao Zhang
Chao Zhang
Citations: 527
h-index: 14
Yixuan Li
Yixuan Li
Citations: 161
h-index: 7
Guangzhi Sun
Guangzhi Sun
Citations: 171
h-index: 2
Yudong Yang
Yudong Yang
Citations: 219
h-index: 8
Wei Li
Wei Li
Citations: 97
h-index: 3
Ma Zejun
Ma Zejun
Citations: 0
h-index: 0

비디오 대규모 언어 모델(LLM)은 컴퓨팅 및 메모리 제약으로 인해 종종 낮은 프레임 속도와 공간 해상도를 사용하는데, 이는 질문 답변(QA)에 필요한 중요한 정보를 놓치게 할 수 있습니다. 실용적이고 효율적인 해결책은 두 단계의 방식으로, 먼저 비디오 전반에 대한 대략적인 이해를 통해 관련 세그먼트를 찾고, 그런 다음 이러한 세그먼트를 더 높은 시간 또는 공간 해상도로 다시 시청하는 것입니다. 본 논문에서는 강화 학습을 통해 반복 시청 기능을 제공하며 체인 오브 씽크(CoT) 초기 설정에 의존하지 않는 최초의 엔드투엔드 비디오 LLM인 video-SALMONN-R$^3$을 소개합니다. 이 설계는 비용이 많이 드는 CoT 데이터 주석의 필요성을 없애고, 사전 학습된 비디오 이해 능력을 저하시킬 수 있는 CoT 기반의 지도 미세 조정(SFT)을 피할 수 있습니다. 반복 시청으로 인해 발생하는 추론 우선 방식과 사전 학습된 비디오 LLM의 답변 우선 경향 사이의 불일치를 해결하기 위해, 모델이 먼저 첫 번째 시청에서 직접적인 답변을 생성하고, 그 후 재시청을 통해 이를 개선하는 재답변 전략을 제안합니다. 마지막으로, 재시청 과정에서 질문 준수도를 향상시키기 위해, 찾은 세그먼트를 다시 검토할 때 질문을 다시 주입하는 재질문 메커니즘을 제안합니다. 실험 결과는 video-SALMONN-R$^3$이 기본 모델과 QA-SFT 기준 모두를 능가하며, 상당한 수준으로 낮은 계산 비용으로 기존의 반복 시청 기반 접근 방식을 뛰어넘는다는 것을 보여줍니다. 코드, 모델 및 데이터는 논문 게재 확정 후 공개될 예정입니다.

Original Abstract

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA). A practical and efficient solution is a two-stage paradigm: first perform coarse video understanding to localize relevant segments, and then re-watch these segments at higher temporal or spatial fidelity. In this paper, we present video-SALMONN-R$^3$, the first end-to-end video-LLM that enables re-watch through reinforcement learning without relying on chain-of-thought (CoT) cold-start. This design removes the need for costly CoT data annotations and avoids CoT-based supervised fine-tuning (SFT), which can otherwise degrade the pretrained video understanding abilities. To address the mismatch between the reasoning-first behavior induced by re-watch and the answer-first tendency of pretrained video-LLMs, we propose a re-answer strategy, in which the model first produces a direct answer in the first watch and then refines it after re-watching. Finally, to improve question adherence during re-watching, we propose a re-ask mechanism that re-injects the query when revisiting localized segments. Experimental results show that video-SALMONN-R$^3$ consistently outperforms both the base model and the QA-SFT baseline, while surpassing prior re-watch-based approaches with significantly lower computational cost. Code, models, and data will be publicly released upon acceptance.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!