2607.17994v1 Jul 20, 2026 cs.CV

HAS: 하이라이트 기반 어텐션 스티어링을 이용한 멀티모달 LLM 비디오 요약

HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization

Yingjie Lao
Yingjie Lao
Citations: 21
h-index: 2
Rui Chu
Rui Chu
Citations: 22
h-index: 1

인공지능(AI) 기술의 발전과 함께 비디오 이해의 중요성이 더욱 커지고 있습니다. 최근에는 멀티모달 대규모 언어 모델(M-LLM)이 비디오 이해 능력을 보여주고 있으며, 효율적인 탐색 및 검색을 위한 중요한 분야인 비디오 요약은 이러한 비디오 이해의 핵심입니다. 비디오 이해와 비디오 요약 모두에서 비디오 내 주요 프레임의 적절한 선택이 필수적입니다. 기존의 비디오 요약 방법들은 주로 선택된 주요 프레임과 관련된 세그먼트 캡션을 중시합니다. 그러나 기존 접근 방식은 전체적인 관점에서 프레임의 중요성을 고려하지 못한다는 한계가 있습니다. 본 논문에서는, 프레임을 개별적으로 선택하여 요약을 수행하는 것이 이해의 일관성을 저해하고 중요한 정보를 누락시키며, MLLM의 잠재력을 낭비할 수 있다는 점을 지적합니다. 따라서, 본 연구에서는 하이라이트 기반 어텐션 스티어링 방법인 HAS를 제안합니다. HAS는 MLLM에게 제공되는 비디오가 연속적인 형태이지만, 동시에 하이라이트 정보를 활용하는 방식을 고려합니다. HAS는 크게 두 부분으로 구성됩니다. 첫 번째 부분은 비디오 전체에 걸쳐 프레임 단위의 연속적인 하이라이트 분포를 찾는 단계입니다. 두 번째 부분은 이러한 하이라이트 분포를 MLLM의 어텐션 스티어링 벡터로 활용하여 비디오 이해도를 향상시키는 단계입니다. 이를 통해 모델 추론 과정에서 하이라이트가 강조된 프레임에 더 많은 집중을 하고, 상대적으로 중요도가 낮은 프레임을 완전히 무시하는 대신 주의를 덜 기울여 전체 정보를 놓치지 않도록 합니다. 우리는 다양한 벤치마크 데이터셋에서 HAS의 성능을 평가했으며, 비디오 요약 분야에서 뛰어난 결과를 보여주었습니다.

Original Abstract

Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation. Recently, Multimodal Large Language Model(M-LLM) has shown its capability in video understanding. Video summarization, a specific domain of video understanding, has proven its importance for efficient navigation and retrieval. Both video understanding and video summarization require a good selection of key frames in a video. Current video summarization methods heavily focus on the selected key frames and correlated segment captions. However, existing approaches overlook the perspective of treating the importance of the frames globally. We argue that using discrete selected frames for summarization will not only reduce the understanding coherence, but also lost important information in the video, as well as wasting the original capacity of the MLLMs. In this paper, we propose HAS, a Highlight-guided Attention Steering method for video summarization. We consider a challenging but practical setting where the video given to MLLMs for summarize should be continuous but with highlight guidance. HAS mainly consists of two parts: The first part is to find a continuous frame-level highlight distribution for the video globally. The second part is to apply the highlight distribution as an attention steering vector for the MLLM, targeting a better understanding of the video, and thus during the model inference time, putting more attention on the highlighted frames, while avoiding lost entire information on less highlighted frames through putting less attention instead of forgetting them. We evaluated HAS on a variety of benchmarks, and it has shown convincing performance in video summarization.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!