2607.24999v1 Jul 27, 2026 cs.CL

CogArena: 대규모 언어 모델의 인지 능력 구조에 대한 다중 방법 평가

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

Kazunori D. Yamada
Kazunori D. Yamada
Citations: 178
h-index: 6
Dengzhe Hou
Dengzhe Hou
Citations: 14
h-index: 2
Lingyu Jiang
Lingyu Jiang
Citations: 11
h-index: 2
Fangzhou Lin
Fangzhou Lin
Citations: 274
h-index: 9

대규모 언어 모델(LLM)의 인지 능력을 평가하는 방식이 점점 더 세분화된 프로필로 요약되고 있으며, 이러한 프로필의 차원은 다양한 작업에서 일관성을 보여야 하고, 특정 개입에 선택적으로 반응해야 하며, 프로필을 정의하는 데 사용된 모델을 넘어 일반화되어야 합니다. 본 연구에서는 다중 방법 프레임워크를 기반으로 13가지 패러다임을 활용한 절차적 생성 벤치마크인 CogArena를 소개합니다. 이를 통해 인지-작업 점수가 이론적으로 뒷받침되는 다섯 가지 그룹으로 분류될 때, 어떤 기준으로 차원 레이블을 부여할 수 있는지 판단합니다. 55개의 공개 모델을 대상으로 분석한 결과, 거의 모든 패러다임 간의 상관관계가 양수이며, 공통적인 요인이 전체 분산의 약 절반을 설명합니다. 그룹 내 성능 우위는 미미하며, 점수에 민감하고 모델 계열에 따라 다릅니다. 12개의 모델(6개 계열)을 대상으로 독립적으로 수행된 교차 연구에서는 특정 유형의 개입이 작은 수준의 그룹 간 차이를 나타내지만, 특정 개입 방식에 따른 효과는 다중 검정으로 인해 유의미하지 않으며, 예측 정확도를 높이지 못합니다. 또한, 미리 정해진 확인 기준은 충족되지 않았습니다. 사후 분석을 통해 표현 방식을 변경한 재현 실험에서는 더 작은 양수 값을 얻었지만, 여전히 실패했습니다. 종합적으로 볼 때, 이러한 결과는 제한적인 결론을 뒷받침합니다. 이론에 기반한 프롬프팅 방식은 주어진 테스트 세트 내에서 약간의 경향성을 보이지만, 현재 증거로는 안정적인 5차원 프로필을 확립하기 어렵습니다. CogArena는 모델 점수에 인지 레이블을 부여하기 전에 행동 특성, 공분산, 맞춤형 개입 및 계열 외부 예측을 결합하는 워크플로우를 제공합니다.

Original Abstract

LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!