뇌 기반 언어 모델을 활용한 견고한 추론을 위한 표현 정렬의 한계
Beyond representational alignment with brain-guided language models for robust reasoning
대규모 언어 모델(LLM)과 인간 고등 인지 기능을 뒷받침하는 신경 메커니즘 간의 상관관계는 아직 충분히 규명되지 않았습니다. 인간 뇌에서 언어와 추론이 분리되는 것처럼 보이는 점을 고려할 때, LLM이 추론 관련 영역의 신경 신호와 일치하는지 여부, 그리고 그러한 신호가 LLM을 향상시킬 수 있는지 여부는 중요한 질문입니다. 본 연구에서는 귀납적 추론에 초점을 맞춰 LLM 내부 표현이 작업 관련 fMRI 활동과 부분적으로 일치할 뿐만 아니라 이러한 신호를 통해 직접 개선될 수 있음을 보여줍니다. 신경 예측 지표를 사용하여 분석한 결과, LLM은 집계 수준에서 추론 관련 영역의 설명 가능한 변동성의 상당 부분을 설명하지만, 특정 유형의 추론에서는 예측력이 낮아 정렬과 불일치를 동시에 나타냅니다. 이러한 결과를 바탕으로 뇌 기반 프레임워크를 제안합니다. 이 프레임워크는 모델과 뇌 표현의 결합 구조에 의해 유도되는 방향으로 모델 표현을 조정하며, 추론 과정에서 개입을 수행하고 학습 과정에서 미세 조정을 적용합니다. 실험 결과, 작업 유발 뇌 신호가 LLM의 추론 능력을 직접적으로 향상시켜 언어 데이터만 사용한 지도 학습 방식과는 다른 방식으로 10개의 LLM(15억~720억 파라미터)에서 성능 향상을 가져왔으며, 다양한 유형의 추론에 대한 일반화 능력과 최대 13%의 절대 정확도 향상을 달성했습니다. 본 연구는 LLM-뇌 간의 상관관계를 단순한 상관관계에서 가이드 중심으로 발전시켜, 보다 견고하고 인지적으로 일치하는 AI 시스템 개발을 위한 뇌 신호 기반 경로를 제시합니다.
The correspondence between large language models (LLMs) and the neural mechanisms underlying human higher-order cognition remains insufficiently characterized. Given that language and reasoning in the human brain appear dissociable, an open question is whether LLMs align with neural signals from reasoning-related regions and whether such signals can improve them. Here, focusing on deductive reasoning, we show that LLM internal representations are not only partially aligned with task-fMRI activity but can also be directly enhanced by these signals. Using a neural-predictivity metric, we find that LLMs explain a substantial fraction of the explainable variance in reasoning-related regions at the aggregate level, whereas predictivity within specific reasoning types is lower, indicating both alignment and divergence. Building on this, we propose a brain-guided framework: we steer model representations along directions induced by the joint structure of model and brain representations, applying intervention at inference and fine-tuning during training. We demonstrate that task-evoked brain signals can directly enhance LLM reasoning, yielding gains orthogonal to language-only supervision across 10 LLMs (1.5B-72B), with transfer across reasoning types and up to 13\% absolute accuracy gain. Our results advance LLM-brain correspondences from correlation to guidance, establishing a brain-signal-driven pathway toward more robust and cognitively aligned AI.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.