언어 이전의 논리: 형식적 유도에 기반한 사전-사전 학습은 기술 습득과 압축성을 향상시킨다
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
기호 데이터로 언어 모델(LM)을 사전-사전 학습하면 자연어 학습 속도를 높이고 성능을 향상시킬 수 있습니다. 그러나 기존의 사전-사전 학습 방식, 예를 들어 Dyck 알고리즘 및 절차적 알고리즘은 자연어의 표현력을 제대로 담아내지 못하는 제한적인 기본 요소에 의존합니다. 또한, 이전 연구들은 상대적으로 작은 토큰 규모로 진행되어 기술 발현과 표현 변화에 대한 제한적인 통찰력만을 제공했습니다. 이러한 한계를 극복하기 위해 우리는 형식적 유도를 활용하여 풍부한 구조적 및 언어적 편향을 부여하는 체계적인 초기화 전략인 논리 사전-사전 학습(Logic-PPT)을 제안합니다. 형식적 유도는 변수 바인딩, 양자 연결, 관계 의존성, 그리고 긴 문맥에서의 술어-논항 구조 결합 등 자연 언어의 핵심 메커니즘을 요구합니다. 100B 토큰 규모로 평가를 확장한 결과, 논리 사전-사전 학습은 LM의 기술 습득 속도를 크게 향상시켰습니다. 표준 초기화 방식보다 36B 토큰 적게 사용하여 80%의 정확도를 달성했으며, 다른 사전-사전 학습 기준 모델보다 우수한 성능을 보였습니다. 메커니즘적으로, 형식적 유도는 지속적인 구조 재구성을 유도하며, 이는 낮은 순위와 스펙트럴 집중성이 특징인 표현 공간으로 나타납니다. 중요한 점은 이러한 내부 기하학이 가지치기를 통해 모델 압축성을 향상시킨다는 것을 보여줍니다. 놀랍게도, 약 33%의 희소성에서도 기존의 밀집 모델과 유사한 성능을 유지할 수 있습니다.
Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To address these limitations, we propose logic pre-pretraining (Logic-PPT) as a principled initialization strategy, leveraging formal derivations to impart richer structural and linguistic biases. Formal derivations require abstract mechanisms that are central to natural language, simultaneously binding variables, connecting quantifiers and relational dependencies, and composing predicate-argument structures over long contexts. Scaling our evaluation to a 100B-token regime, logic pre-pretraining substantially accelerates skill acquisition in LMs, achieving 80\% accuracy on linguistic tasks with 36B fewer tokens than standard initialization, and outperforming alternative pre-pretraining baselines. Mechanistically, formal derivations induce persistent structural reorganization, distinctively characterized by a lower-rank, spectrally concentrated representation space. Crucially, we show that this internal geometry enables improved model compressibility via pruning, matching the dense baseline performance even at $\approx$33\% sparsity.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.