시선-텍스트 생성: 인간의 주의에 대한 범주화된 해독을 넘어
Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
본 논문에서는 새로운 학습 문제를 제안합니다. 이는 다양한 시각적 작업에서 인간의 목표를 자연어 설명으로 변환하는 것입니다. 기존 연구가 시선 해독을 미리 정의된 범주에 대한 판별적인 작업으로 간주하는 반면, 우리는 이를 생성적 학습 문제로 규정하여 모델이 자유 형식의 설명을 생성하도록 훈련합니다. 이 설명은 고정된 레이블을 넘어 인간의 의도의 풍부한 미묘함과 개방성을 포착합니다. 이러한 목표를 달성하기 위해 시선-텍스트 변환 프레임워크인 Gazette를 소개합니다. Gazette는 다중 모드 대규모 언어 모델(MLLM)을 기반으로 하며, 시선 스캔 경로를 자연어로 해독하여 범주화된 레이블을 넘어 인간의 목표를 설명하는 데 사용됩니다. Gazette가 개별적인 시선 행동의 차이를 필터링하고 정확한 자연어 목표 설명을 생성하는 데 중요한 목표 특정 시공간 역학을 학습할 수 있도록, 대규모 언어 모델의 방대한 지식과 추론 능력을 활용하여 목표 지향적 주의 행동에 대한 자연어 설명을 생성하는 '생각-말하기' 기록이라는 새로운 전략을 제안합니다. 이러한 합성 내러티브를 사용한 명령어 튜닝을 통해 Gazette는 여러 작업에서 시선 해독 성능이 최고 수준을 달성하며, 일반화 능력과 다양성을 입증합니다. 이를 통해 시선은 다양한 상황에서 인간의 목표와 의도를 추론하는 강력하고 비침투적인 단서로 활용될 수 있습니다.
We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as a discriminative task over predefined categories, we formulate it as a generative learning problem: training a model to produce free-form descriptions that capture the rich nuances and open-ended nature of human intentions beyond fixed labels. To this end, we introduce Gazette, the first gaze-to-text decoding framework. Based on multimodal large language models (MLLMs), Gazette learns to decode gaze scanpaths into natural language for goals that may extend beyond categorical labels and require articulation in natural language. To help Gazette filter out individual differences in gaze behavior and learn the goal-specific spatiotemporal dynamics crucial for generating accurate natural language goal descriptions, we propose a novel strategy that leverages the encyclopedic knowledge and reasoning abilities of a large language model to synthesize natural language explanations of goal-directed attentional behavior called think-aloud transcripts. Instruction tuning on these synthetic narratives allows Gazette to achieve state-of-the-art performance in gaze decoding across multiple tasks, demonstrating its generalizability and versatility, thereby enabling gaze to serve as a powerful, non-intrusive cue for inferring human goals and intentions in diverse scenarios.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.