파괴된 구성성: 산술 연산을 위한 트랜스포머의 직관에 어긋나는 학습 동역학
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
대규모 언어 모델(LLM)은 규모가 커져도 예상치 못한 오류나 의도하지 않은 동작을 보이는 경우가 많습니다. 최근 연구에서는 LLM과 인간의 기술 조합 능력 간의 차이가 밝혀졌지만, 기술 조합의 학습 동역학과 인간과 다른 행동의 근본적인 원인은 여전히 명확하지 않습니다. 본 연구에서는 트랜스포머 모델을 인공적인 산술 작업에 훈련시켜 학습 동역학의 메커니즘을 조사합니다. 광범위한 분석과 세분화된 진단 지표를 통해, 트랜스포머 모델이 인간과 유사한 순차적 규칙에 따라 안정적으로 기술 조합을 구성하지 못한다는 사실을 발견했습니다. 오히려 모델은 종종 기술을 역순으로 또는 병렬적으로 학습하며, 이는 데이터 분포 변화 하에서 예상치 못한 혼합 오류를 야기하는데, 이러한 현상을 우리는 '파괴된 구성성(shattered compositionality)'이라고 부릅니다. 이러한 현상을 설명하기 위해, 모델이 원인-결과 관계나 절차적 구성이 아닌, 훈련 데이터와의 상관 관계에 기반하여 학습 동역학을 형성한다는 증거를 제시합니다. 또한, '파괴된 구성성'이 최신 LLM에서도 지속적으로 나타나며, 단순히 모델 크기를 키우거나 스크래치패드 기반 추론을 적용해도 완화되지 않는다는 것을 보여줍니다. 본 연구의 결과는 모델의 학습 행동과 바람직한 기술 조합 간의 근본적인 불일치를 드러내며, 이는 추론의 신뢰성, 데이터 분포 변화에 대한 강건성, 그리고 모델 정렬에 중요한 영향을 미칩니다.
Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanistic studies reveal the discrepancy between LLMs and humans in skill compositions, the learning dynamics of skill acquisition and the role of data distributions remain elusive. In this study, we train transformers on synthetic arithmetic tasks with black-box model-agnostic metrics for analyzing non-human skill compositions. We discover that transformers often acquire skills for arithmetic in reverse order or in parallel instead of human-like sequential rules--a phenomenon we refer to as shattered compositionality. To explain these behaviors, we provide evidence that correlational matching to the training data, rather than causal or procedural composition, shapes learning dynamics. As a consequence, this non-human acquisition creates competition between partially learned skills, producing characteristic mixing errors and weaker robustness under controlled distribution shifts. We further show that the same qualitative behavior persists in modern LLMs and is not mitigated by pure model scaling or scratchpad supervision. Our results highlight a mismatch between training-time skill acquisition and the human-like hierarchical compositions, with implications for reasoning reliability and out-of-distribution robustness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.