NeuroCogMap: 거대 언어 모델의 인지적 구조를 밝히다
NeuroCogMap Reveals Cognitive Organization of Large Language Models
인공 시스템 내에서 복잡한 인지 기능이 어떻게 조직되어 있는지 이해하는 것은, 거대 언어 모델(LLM)을 해석하고 이를 생물학적 인지와 연관시키는 데 매우 중요합니다. 그러나 LLM은 다양한 인지적인 행동을 보이지만, 그 내부 표현이 실제로 행동, 실패 및 인간 인지와의 연결을 설명하는 재현 가능한 기능 시스템을 형성하는지에 대해서는 명확하지 않습니다. 본 연구에서는 NeuroCogMap이라는 인지 신경과학에 영감을 받은 프레임워크를 제시합니다. 이 프레임워크는 LLM의 내부 특징들을 기능적 영역으로 나누고, 이를 해석 가능한 기능, 인지 능력 및 인지 계층 구조와 연결합니다. 이러한 영역들은 안정적이고 의미적으로 일관된 조직을 형성하며, 이는 모델 간에 부분적으로 보존되고 모델 출력과 기능적으로 연관되어 있습니다. 이 조직 내에서 LLM의 주요 실패 사례(예: 환각, 편향, 거부 실패, 아첨)는 표현 및 행동 제어 시스템의 뚜렷한 오류와 일치하며, 이를 통해 메커니즘 기반 감지 및 표적 개입을 위한 내부적인 특징을 제공합니다. 모델의 행동 범위를 넘어, NeuroCogMap은 자연스러운 언어 이해 과정에서 인간 피질 반응을 예측하는 데 기여하며, 특히 고차 연합 피질 영역에서 가장 강력한 상관관계를 보입니다. 인지 수준에서는, 이 프레임워크의 내부 특징들이 인간 의사 결정에 대한 기존 모델 개선을 위한 잠재적인 전략을 드러냅니다. 종합적으로, 이러한 결과들은 NeuroCogMap이 인공 시스템에서의 기능적 조직을 매핑하고, 이를 인간 피질 기능 및 인지 행동과 연관시키는 시스템 수준의 프레임워크임을 입증합니다.
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.