2605.29588v1 May 28, 2026 cs.CV

Brain-IT-VQA: 뇌 신호를 활용한 시각 질의 응답

Brain-IT-VQA: From Brain Signals to Answers

Roman Beliy
Roman Beliy
Citations: 307
h-index: 5
Matias Cosarinsky
Matias Cosarinsky
Citations: 6
h-index: 1
Oliver Heinimann
Oliver Heinimann
Citations: 12
h-index: 2
Navve Wasserman
Navve Wasserman
Citations: 121
h-index: 5
Michal Irani
Michal Irani
Citations: 52
h-index: 5

사람이 이미지를 보면서 기록되는 fMRI 신호로부터 시각적 내용을 해독하고, 특히 해당 이미지에 대한 질문에 답하는 것은 오랜 과제입니다. 최근 몇 년 동안 fMRI를 이용한 시각 질의 응답(VQA) 분야에서 상당한 발전이 있었지만, 성능은 여전히 제한적입니다. 또한, 최신 모델들이 점점 더 정확한 예측을 수행할 수 있게 되었음에도 불구하고, 이러한 모델들은 뇌 내 시각적 표현 구조를 이해하는 도구로 거의 사용되지 않았습니다. 본 논문에서는 fMRI로부터 시각 질의 응답을 위한 프레임워크인 Brain-IT-VQA를 제시합니다. Brain Interaction Transformer (Brain-IT)를 기반으로, 저희 방법은 뇌 활동으로부터 언어 토큰을 해독하고 이를 언어 모델과 통합하여 시각적 질문에 답변합니다. 저희 모델은 기존의 fMRI 기반 이미지 설명 및 VQA 접근 방식보다 훨씬 뛰어난 성능을 보입니다. 또한, 저희는 fMRI로부터 시각 질의 응답을 위한 새로운 데이터셋 및 벤치마크인 NSD-VQA를 소개합니다. 기존의 이미지-fMRI VQA 데이터셋은 일반적으로 각 이미지에 대해 몇 개의 광범위하고 약하게 제어된 질문만 제공하는 반면, NSD-VQA는 평균적으로 각 이미지당 20개의 질문-답변 쌍을 제공하며, 이는 여러 수준의 시각적 이해를 분리하는 20가지의 제어된 질문 범주로 구성되어 있습니다. 이를 통해 제한적인 fMRI 테스트 데이터에도 불구하고 더욱 신뢰성 있고 해석 가능한 평가가 가능합니다. Brain-IT-VQA와 NSD-VQA는 강력한 예측 프레임워크뿐만 아니라 뇌 표현을 연구하기 위한 도구도 제공합니다. 이 벤치마크를 사용하여 자연 이미지에 대한 fMRI 반응으로부터 안정적으로 해독할 수 있는 시각 및 의미 정보의 형태를 정량화했습니다. 또한, 다양한 질문 유형에 따른 서로 다른 뇌 영역의 기여도를 분석했습니다.

Original Abstract

Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While significant progress has been made in recent years in visual question answering (VQA) from fMRI, performance remains limited. Moreover, although recent models can make increasingly accurate predictions, they have rarely been used as tools for understanding the structure of visual representations in the brain. We present Brain-IT-VQA, a framework for visual question answering from fMRI. Building on the Brain Interaction Transformer (Brain-IT), our method decodes language tokens from brain activity and integrates them with a language model to answer visual questions. Our model substantially outperforms previous fMRI-based captioning and VQA approaches. We further introduce NSD-VQA, a new dataset and benchmark for visual question answering from fMRI. Unlike existing image-fMRI VQA datasets, which typically provide only a few broad and weakly controlled questions per image, NSD-VQA provides on average 20 question-answer pairs per image across 20 controlled question categories that disentangle multiple levels of visual understanding. This enables more reliable and interpretable evaluation despite limited fMRI test data. Together, Brain-IT-VQA and NSD-VQA provide both a strong predictive framework and a tool for studying brain representations. Using this benchmark, we quantify which forms of visual and semantic information can be reliably decoded from fMRI responses to natural images. We further analyze the contributions of different brain regions across question types.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!