시를 통해 시인의 출신 지역을 예측하는 연구: 완전 당나라 시집에 나타난 지역별 언어적 특징에 대한 계산 분석
Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems
본 연구는 당나라 시대 시인들의 작품에서 그들의 지리적 출신이 감지 가능한 언어적 흔적을 남기는지 질문합니다. '완전 당나라 시집(Quan Tang Shi)'에 기재된 모든 시를 각 저자별로 분류하고, 중국 전기 데이터베이스(CBDB)를 통해 시인들을 해당 지역의 행정구역과 연결하여, 10개의 당나라 행정구역에 걸쳐 357명의 시인을 포함하는 시인 수준의 코퍼스를 구축했습니다. 본 연구에서는 출신지 예측을 다중 분류 문제로 정의하고, 문자 n-그램 TF-IDF와 함께 해석 가능한 도메인 특징(이미지, 계절, 은유)을 사용하여 고전 모델과 신경망 모델을 통해 시인의 광범위한 지역(남부 vs. 북부)을 69%의 정확도로 예측했으며, 이는 단순 다수결 기준선인 53%보다 훨씬 높은 수치입니다. 또한 행정구역 수준의 출신지를 우연에 의해 가능한 것 이상으로 정확하게 예측했습니다. 분류 분석 외에도 세 가지 중요한 결과가 도출되었습니다. (i) 행정구역 간의 언어적 거리는 지리적 거리와 함께 증가합니다 (Mantel r = 0.40, p ≈ 0.09, 9개의 행정구역 기준). 이는 시적 언어에서 거리 감소 효과를 보여주는 증거입니다. (ii) 이 현상은 시간에 따라 변화하는데, 고당 시대에는 남부와 북부의 구분이 우연에 가깝지만, 후기 당나라 시대에는 가장 두드러집니다. 이는 제국 전성기에 궁정 주도의 획일화가 이루어졌다가 이후 지역 간의 다양성이 증가했음을 시사합니다. (iii) 모델의 확신을 가진 오답은 역사적으로 의미 있는 정보를 담고 있습니다. 초기 당나라 시대에는 모든 오분류가 남부 출신의 시인을 북부 출신으로 잘못 분류한 경우였으며, 이는 북부 궁정의 언어 사용에 대한 권위를 반영합니다. 또한 계층적 frozen-encoder 표현을 사용하여 전체 코퍼스를 분석했을 때, 고전 중국어 변환기 모델(GuwenBERT)은 단순 TF-IDF보다 성능이 더 뛰지 못했으며, 두 가지 방법을 결합하는 것은 추가적인 이점을 제공하지 않았습니다. 이는 문자 n-그램이 이미 지역적 특징을 잘 포착하고 있음을 나타냅니다. 본 연구의 결과는 해석 가능한 머신러닝을 문학사 연구를 위한 가설 생성 도구로 활용할 수 있음을 보여줍니다.
We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work. Aggregating every poem attributed to each author in the Complete Tang Poems (Quan Tang Shi) and linking poets to their administrative circuit of origin via the China Biographical Database (CBDB), we build a poet-level corpus of 357 poets across the ten Tang circuits and frame origin prediction as multi-class classification. Using character $n$-gram TF-IDF together with interpretable domain features (imagery, season, and allusion), classical and neural models predict a poet's broad region (South vs.\ North) at $0.69$ accuracy, well above the $0.53$ majority baseline, and finer circuit-level origin above chance. Beyond classification, three findings emerge. (i) Linguistic distance between circuits grows with geographic distance (Mantel $r=0.40$, $p\approx0.09$ over nine circuits), evidence of a distance-decay effect in poetic language. (ii) The signal interacts with time: South/North separability is at chance in the High Tang and strongest in the Late Tang, consistent with court-driven homogenization at the empire's height followed by regional divergence. (iii) The model's confident errors are historically meaningful -- in the Early Tang, every misclassification is a southern poet read as northern, reflecting the prestige of the northern court idiom. We further show that, when given the whole corpus through a hierarchical frozen-encoder representation, a classical-Chinese transformer (GuwenBERT) only matches -- not beats -- simple TF-IDF, and that combining them adds nothing, indicating that character $n$-grams already capture the regional signal. Our results position interpretable machine learning as a hypothesis generator for literary history.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.