언어의 문맥적 표현에서 나타나는 난류와 유사한 5/3 스펙트럼 스케일링: 복잡계로서의 언어
Turbulence-like 5/3 spectral scaling in contextual representations of language as a complex system
자연어는 강력한 통계적 규칙을 보이는 복잡계입니다. 본 연구에서는 텍스트를 트랜스포머 기반 언어 모델에 의해 생성된 고차원 임베딩 공간에서의 경로로 표현하고, 토큰 시퀀스에 따른 스케일 의존적 변동성을 임베딩-스텝 신호를 사용하여 정량화합니다. 여러 언어 및 코퍼스에서 얻은 결과, 생성된 파워 스펙트럼은 넓은 주파수 범위에 걸쳐 약 5/3의 지수를 갖는 견고한 거듭제곱 법칙을 나타냅니다. 이러한 스케일링은 인간이 작성한 텍스트와 AI가 생성한 텍스트 모두에서 얻은 문맥적 임베딩에서 일관되게 관찰되지만, 정적 단어 임베딩에서는 나타나지 않으며, 토큰 순서의 무작위화에 의해 파괴됩니다. 이러한 결과는 관찰된 스케일링이 어휘 통계뿐만 아니라 다중 스케일, 문맥 의존적인 조직을 반영한다는 것을 보여줍니다. 난류에서의 콜모고로프 스펙트럼과의 유사성을 통해, 본 연구의 결과는 의미 정보가 언어적 스케일 전반에 걸쳐 스케일-프리(scale-free)하고 자기-유사(self-similar)한 방식으로 통합된다는 것을 시사하며, 언어 표현의 복잡한 구조를 연구하기 위한 정량적이고 모델에 독립적인 벤치마크를 제공합니다.
Natural language is a complex system that exhibits robust statistical regularities. Here, we represent text as a trajectory in a high-dimensional embedding space generated by transformer-based language models, and quantify scale-dependent fluctuations along the token sequence using an embedding-step signal. Across multiple languages and corpora, the resulting power spectrum exhibits a robust power law with an exponent close to $5/3$ over an extended frequency range. This scaling is observed consistently in contextual embeddings from both human-written and AI-generated text, but is absent in static word embeddings and is disrupted by randomization of token order. These results show that the observed scaling reflects multiscale, context-dependent organization rather than lexical statistics alone. By analogy with the Kolmogorov spectrum in turbulence, our findings suggest that semantic information is integrated in a scale-free, self-similar manner across linguistic scales, and provide a quantitative, model-agnostic benchmark for studying complex structure in language representations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.