2605.26789v1 May 26, 2026 cs.AI

구성적 오류: 안정적인 사실 지식이 반드시 구성적 추론 능력을 의미하지 않는다

Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning

Hongzhi Wang
Hongzhi Wang
Citations: 6
h-index: 2
Wenpeng Xing
Wenpeng Xing
Citations: 205
h-index: 10
Zhengtao Yu
Zhengtao Yu
Citations: 20
h-index: 2
Xuyang Teng
Xuyang Teng
Citations: 6
h-index: 2
Meng Han
Meng Han
Citations: 78
h-index: 3
Yunzhao Wei
Yunzhao Wei
Citations: 0
h-index: 0
Jiefeng Chen
Jiefeng Chen
Citations: 150
h-index: 4

사후 학습 모델의 성능 평가는 일반적으로 여러 단계의 추론을 하나의 능력으로 간주하는 종합적인 벤치마크 점수를 통해 이루어집니다. 즉, 더 많은 질문에 정확하게 답변하는 모델이 반드시 사실 정보를 효과적으로 조합하는 능력이 뛰어나다고 가정합니다. 본 연구에서는 이러한 가정이 오해를 불러일으킬 수 있음을 보여줍니다. 통계적으로 차이가 없는 기본 지식을 가진 레시피들이 40% 이상의 격차를 보이는 구성적 행동을 나타내는 현상, 즉 '구성적 오류'가 발생하며, 이는 종합적인 지표로는 감지하기 어렵습니다. 우리는 새로운 평가 프로토콜인 '이중 게이트' 방식을 도입하여, 종합적인 구성 능력의 차이를 넘어 안정적으로 접근 가능한 사실 정보에 조건부로 발생하는 잔여 구성 실패를 측정합니다. 이를 통해 사후 학습으로 얻은 성능 향상을 세 가지 독립적인 요소, 즉 기본 지식의 안정성, 잔여 구성 능력 및 핵심 깊이로 분해했습니다. 4가지 사후 학습 레시피를 사용한 시간적 사실 정보 연쇄 벤치마크에서 이러한 분석 결과는 사후 학습 목표가 종합적인 지표로는 파악하기 어려운 방향으로 구성 능력을 변화시킨다는 것을 보여주며, 다단계 추론 능력 향상에 대한 주장은 반드시 기본 지식 접근성을 고려한 세분화된 평가 지표와 함께 제시되어야 함을 시사합니다. 추가적으로 수행한 진단 분석 결과, 측정된 구성 실패의 상당 부분은 모델이 구성 능력이 부족해서 발생하는 것이 아니라, 생성 시간 동안 발생하는 계산 제약으로 인해 발생하는 현상임을 확인했습니다.

Original Abstract

Post-training is routinely evaluated through aggregate benchmark scores that treat multi-hop reasoning as a single capability -- as if a model that answers more questions correctly must be better at assembling facts. We show that this assumption can be misleading: recipes with statistically indistinguishable atomic knowledge produce composition behaviour separated by over 40 percentage points, a phenomenon we call composition collapse: the systematic failure to assemble stably-known facts into chains, invisible to aggregate metrics. We introduce a double-gate protocol that changes the estimand from an aggregate compositionality gap to residual composition failure conditioned on stable atomic access, decomposing post-training gains into three independent channels: atomic stability, residual composition, and critical depth. On a benchmark of temporal factual chains spanning depths 2--11 across four post-training recipes, this decomposition reveals that post-training objectives shift composition capability in directions that aggregate metrics mask, and suggests that claims about multi-hop reasoning improvement should be accompanied by atomic-gate-controlled composition metrics. Diagnostic probes further show that a substantial share of measured composition failure reflects generation-time computation constraints rather than permanent inability to compose.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!