2608.01980v1 Aug 03, 2026 cs.CV

AdaThinkV: 토큰 효율적인 비디오 추론을 위한 적응적 사고

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

Hongbo Jin
Hongbo Jin
Citations: 36
h-index: 4
Shilin Ma
Shilin Ma
Citations: 24
h-index: 4
Jingqi Tian
Jingqi Tian
Citations: 36
h-index: 3
Haoji Zhang
Haoji Zhang
Citations: 408
h-index: 7
Lin Chen
Lin Chen
Citations: 0
h-index: 0
Haonan Xu
Haonan Xu
Citations: 0
h-index: 0
Tianrui Zhu
Tianrui Zhu
Citations: 49
h-index: 3
Xin Shui
Xin Shui
Citations: 97
h-index: 6
Wenjing Yang
Wenjing Yang
Citations: 1
h-index: 1
Yansong Tang
Yansong Tang
Citations: 124
h-index: 5

사고 과정(Chain-of-thought, CoT) 추론은 어려운 비디오 질문에 대한 성능을 향상시킬 수 있지만, 종종 단순한 질문에는 불필요하게 많은 디코딩 토큰을 소비합니다. 본 연구에서는 멀티모달 대규모 언어 모델이 각 질문에 대해 자신의 추론 노력을 어떻게 조정할 수 있는지 조사합니다. 우리는 오프라인 난이도 레이블, 수동으로 튜닝된 신뢰도 임계값 또는 외부 라우터 없이 명시적인 추론을 수행할지 여부를 학습하는 적응형 프레임워크인 AdaThinkV를 제안합니다. 강화 학습 과정에서 AdaThinkV는 각 프롬프트에 대해 명시적 추론과 직접 답변 모드 간의 일치된 시뮬레이션을 샘플링합니다. ThinkGain은 명시적 추론의 정확도 향상과 추가적인 응답 길이를 균형 있게 고려하여 프롬프트 수준에서의 유용성을 추정하고, 조건부 응답 생성 및 자율 모드 선택에 대한 지침을 제공합니다. 어려운 프롬프트의 경우, 제한된 시뮬레이션 탐색은 모든 응답이 실패하고 정확도 보상이 거의 변동하지 않는 그룹을 야기하여 학습에 충분한 정보를 제공하지 못할 수 있습니다. 따라서 우리는 정보적인 신호를 회복하기 위해 이러한 그룹을 유지하고 점진적으로 확장하는 Variance Recovery Policy Optimization (VRPO)를 도입합니다. 추론 단계에서 AdaThinkV는 응답 모드를 선택하고 단일 자동 회귀 시퀀스로 응답을 생성합니다. 통일된 비디오 추론 평가 세트 전반에 걸쳐 AdaThinkV는 평균 정확도 40.79%를 달성했으며, 평균 출력 토큰 수는 257.20개로, 가장 강력한 적응형 기준 모델보다 2.98 포인트 더 높은 성능을 보였으며, 동시에 22.7% 더 적은 토큰을 사용했습니다. 프로젝트 페이지: https://trilarflagz.github.io/AdaThinkV/

Original Abstract

Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large language model can adapt its reasoning effort to each question. We propose AdaThinkV, an adaptive framework for video reasoning that learns whether to reason explicitly without offline difficulty labels, manually tuned confidence thresholds, or an external router. During reinforcement learning, AdaThinkV samples matched rollouts in explicit reasoning and direct answering modes for each prompt. ThinkGain estimates the prompt-level utility of explicit reasoning by balancing its accuracy gain against additional response length, providing supervision for both conditional response generation and autonomous mode selection. For difficult prompts, limited rollout exploration can yield groups in which every response is unsuccessful and accuracy rewards show little variation, providing insufficient signal for learning. We therefore introduce Variance Recovery Policy Optimization (VRPO), which retains and progressively expands these groups to recover informative signals from prompts that are difficult yet solvable. At inference, AdaThinkV selects a response mode and generates the response in a single autoregressive sequence. Across a unified suite of video reasoning evaluations, AdaThinkV achieves a mean accuracy of 40.79 with an average of 257.20 output tokens, outperforming the strongest evaluated adaptive baseline by 2.98 points while using 22.7% fewer tokens. Project page: https://trilarflagz.github.io/AdaThinkV/

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!