ChronoLens: 시간, 언어 및 언어학적 수준에 따른 언어 변화 측정
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
역사적인 언어 변화는 형태론, 구문, 의미론 및 화용론에 영향을 미치지만, 기존의 계산 연구에서는 이러한 수준들이 호환되지 않는 표현 방식을 사용하여 언어 간에 함께 진화하는지 여부를 판단하기 어렵습니다. 우리는 단일 분석 공간 내에서 변화의 크기와 방향이 언어학적 수준, 언어 및 역사적 시기에 따라 어떻게 변하는지를 조사하여 이 문제를 해결하고자 합니다. 우리는 ChronoLens라는 프레임워크를 소개합니다. 이는 동결된 다국어 언어 모델, 특징 정렬 크로스 코더 및 사후 언어학적 개입을 결합하며, 이를 1803년부터 2026년까지의 다섯 가지 의회 기록에서 추출한 4498만 건의 문서와 약 172억 개의 토큰에 적용했습니다. 결과적으로 생성된 희소 표현은 기존의 밀집 임베딩 또는 풀링된 희소 오토인코더보다 언어학적 통계와 훨씬 더 높은 상관관계를 보입니다 (ρ=0.72 대 0.29 및 0.28). 또한, 형태론, 구문, 의미론 및 화용론은 일반적으로 하나의 언어 내에서 비슷한 정도로 변화하지만, 언어들은 언제, 얼마나, 그리고 어떤 방향으로 변화하는지에 있어서 뚜렷한 차이를 보입니다. 이러한 결과는 역사적인 언어 변화가 체계적이고 다차원적인 과정임을 보여줍니다. 유사한 크기의 변화가 서로 다른 경로를 가릴 수 있으며, 의미 있는 교언어 비교를 위해서는 거리와 방향 모두를 측정해야 합니다.
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this problem by asking how the magnitude and direction of change vary across linguistic levels, languages, and historical periods within a single analytical space. We introduce ChronoLens, a framework that combines frozen multilingual language models, feature-aligned crosscoders, and post-hoc linguistic interventions, and apply it to 44.98 million documents and approximately 17.2 billion tokens from five parliamentary traditions spanning 1803--2026. The resulting sparse representations agree substantially more strongly with linguistic statistics than dense embeddings or a pooled sparse autoencoder ($ρ=0.72$ versus $0.29$ and $0.28$), and reveal that morphology, syntax, semantics, and pragmatics generally change by comparable amounts within a language, while languages differ markedly in when, how far, and in which direction they change. These findings show that historical language change is a structured, multidimensional process: similar magnitudes can conceal different trajectories, and meaningful cross-linguistic comparison requires measuring both distance and direction.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.