VEGAS: 시선 정보를 활용한 인간 중심 비디오 자막 평가
VEGAS: Human-Aligned Video Caption Evaluation via Gaze
비전-언어 모델은 비디오 자막 생성에 뛰어난 성능을 보이지만, 일반적으로 개별 시청자의 주의를 제대로 반영하지 못하는 설명을 생성합니다. 본 논문에서는 VEGAS(Video caption Evaluation via GAze Score)라는 새로운 평가 지표를 제안합니다. VEGAS는 학습 과정 없이 테스트 시간에 수집된 시선 정보를 활용하여 개인화되고 주의 집중 영역과 일치하는 텍스트를 추출합니다. 이는 크로스 모달 방식의 정보 이론 기반 지표로서, 후보 자막이 시청자의 관심 영역을 얼마나 잘 반영하는지를 정량적으로 측정합니다. VEGAS의 성능을 평가하기 위해, 우리는 개인 활동 및 교육 슬라이드 데이터셋을 구축하고, 동기화된 시선 정보와 참조 주석을 함께 제공했습니다. 그런 다음, VEGAS를 사용하여 모델 재학습 없이 거부 샘플링 방식으로 자막을 선택했습니다. 실험 결과, VEGAS에 의해 선택된 자막은 인간의 집중 영역과 훨씬 더 잘 일치하며, 다운스트림 비디오-자막 검색 성능을 향상시킵니다. 이는 추론 과정에서 시청자의 주의를 고려하는 것이 실제적으로 유용하다는 것을 보여줍니다.
Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, attention-aligned text. It is a cross-modal, information-theoretic metric that quantifies how well a candidate caption matches a viewer's focus. To evaluate VEGAS, we curate a dataset of egocentric activities and instructional slides paired with synchronized gaze and reference annotations. We then select captions based on VEGAS via rejection sampling without model retraining. Experiments show that VEGAS-selected captions align significantly better with human focus and improve downstream caption-to-video retrieval, demonstrating the practical utility of incorporating viewer attention during inference.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.