2602.13194v2 Feb 13, 2026 cs.CL

의미론적 덩어리 나누기 및 자연어의 엔트로피

Semantic Chunking and the Entropy of Natural Language

Weishun Zhong
Weishun Zhong
Citations: 9
h-index: 2
Doron Sivan
Doron Sivan
Citations: 7
h-index: 2
Tankut Can
Tankut Can
Citations: 34
h-index: 3
M. Katkov
M. Katkov
Citations: 436
h-index: 13
M. Tsodyks
M. Tsodyks
Citations: 16,140
h-index: 53

인쇄 영어의 엔트로피율은 약 1비트/문자 정도로 추정되며, 이는 현대 대규모 언어 모델(LLM)이 최근에야 달성한 수준입니다. 이러한 엔트로피율은 영어가 무작위 텍스트에서 예상되는 5비트/문자 대비 약 80%의 중복성을 가지고 있음을 의미합니다. 본 연구에서는 자연어의 복잡한 다중 스케일 구조를 포착하려는 통계 모델을 제시하며, 이를 통해 이러한 중복성의 수준을 근본적으로 설명하고자 합니다. 제안하는 모델은 텍스트를 의미적으로 일관된 덩어리로 자체 유사하게 분할하는 절차를 설명하며, 이를 통해 텍스트의 의미 구조를 계층적으로 분해하여 분석할 수 있습니다. 현대 LLM 및 공개 데이터 세트를 사용한 수치 실험 결과, 제안하는 모델이 다양한 의미 계층 수준에서 실제 텍스트의 구조를 정량적으로 잘 나타냄을 확인했습니다. 모델이 예측하는 엔트로피율은 인쇄 영어의 추정 엔트로피율과 일치합니다. 또한, 본 연구는 자연어의 엔트로피율이 고정된 값이 아니라, 모델의 유일한 자유 변수에 의해 표현되는 코퍼스의 의미적 복잡성에 따라 체계적으로 증가해야 함을 밝힙니다.

Original Abstract

The entropy rate of printed English is famously estimated to be about one bit per character, a benchmark that modern large language models (LLMs) have only recently approached. This entropy rate implies that English contains nearly 80 percent redundancy relative to the five bits per character expected for random text. We introduce a statistical model that attempts to capture the intricate multi-scale structure of natural language, providing a first-principles account of this redundancy level. Our model describes a procedure of self-similarly segmenting text into semantically coherent chunks down to the single-word level. The semantic structure of the text can then be hierarchically decomposed, allowing for analytical treatment. Numerical experiments with modern LLMs and open datasets suggest that our model quantitatively captures the structure of real texts at different levels of the semantic hierarchy. The entropy rate predicted by our model agrees with the estimated entropy rate of printed English. Moreover, our theory further reveals that the entropy rate of natural language is not fixed but should increase systematically with the semantic complexity of corpora, which are captured by the only free parameter in our model.

0 Citations
0 Influential
26.5 Altmetric
132.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!