2607.27595v1 Jul 30, 2026 cs.CL

유사성 너머: 고전 중국 역사에서 문맥적 연관성을 추출하고 전문가가 평가하는 방법

Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories

Zhaoji Wang
Zhaoji Wang
Citations: 1,504
h-index: 5
Wanyu Si
Wanyu Si
Citations: 1
h-index: 1
Jun Wang
Jun Wang
Citations: 0
h-index: 0

문맥적 연관성에 대한 계산적 접근 방식은 문자열 매칭에서 신경망 기반 검색으로 발전했지만, 이러한 접근 방식의 결과물인 유사도 점수와 병렬 구절 목록은 텍스트들이 서로를 어떻게 재사용하는지 또는 그 이유를 설명하지 못합니다. 본 연구에서는 세밀한 수준의 문맥적 연관성 추출을 일종의 '주체적인' 작업으로 재구성했습니다. 대규모 언어 모델(LLM)이 두 개의 텍스트 단위를 전체적으로 읽고, 제한된 도구 인터페이스를 통해 제안된 각 재사용 사례를 양쪽 텍스트 내의 정확한 문자 범위에 연결하고, 다섯 가지 차원의 유형(형식, 측면, 출처 표시, 기능, 태도) 중 하나로 분류해야 합니다. 본 연구에서는 논어와 한서에 대한 포괄적인 비교 분석을 통해 이 접근 방식을 검증했습니다. 세 명의 분야 전문가가 여러 모델에서 생성된 후보 집합을 평가하여 2,533개의 문맥적 연관성 쌍으로 구성된 표준 데이터셋을 구축했습니다. 이 표준 데이터셋을 기준으로 12개의 LLM을 분석한 결과, 정확도(56%-93%), 유사한 품질 수준에서의 51배에 달하는 비용 차이, 그리고 각 모델의 신뢰도 값이 얼마나 잘 조정되었는지 확인했습니다. 전문가 간의 합의도는 신뢰성 기울기를 보여주었습니다. 즉, 텍스트 표면에서 명확하게 보이는 측면은 일관되게 주석 처리되는 반면, 의도를 추론해야 하는 측면은 논쟁의 여지가 있으며, 이는 이러한 주석이 뒷받침할 수 있는 주장의 범위를 제한합니다. 검증된 추출 도구를 전체 24사(65,380개의 비교 항목, 5,766개의 쌍)에 적용한 결과, 유사도 점수로는 표현할 수 없는 코퍼스 수준의 구조를 파악할 수 있었습니다. 인용문의 해석적 구성은 18세기에 걸쳐 체계적으로 변하지 않았지만, 동일한 구절이 점점 더 간접적인 방식으로 인용되는 경향을 보였습니다. 이는 전체적으로는 안정성을 유지하면서 개별 사례에서는 변화가 나타나는 현상으로, 문화적 매력 이론에서 예상되는 결과입니다. 본 연구에서는 추출 프로토콜과 전문가 평가를 거친 표준 데이터셋을 공개합니다.

Original Abstract

Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how or why. We recast fine-grained intertextuality extraction as an agentic task in which a large language model (LLM) reads two text units in full and, through a constrained tool interface, must ground each proposed reuse in exact character spans on both sides and label it under a five-dimension typology of reuse (form, aspect, source-marking, function, stance). We validate the approach on an exhaustive comparison of the Analects with the Book of Han, where three domain experts adjudicate a pooled multi-model candidate set into a benchmark of 2,533 intertextual pairs. Against this standard we study twelve LLMs, reporting precision (56%-93%), a 51$\times$ cost spread at comparable quality, and how well their confidence is calibrated. Expert agreement traces a reliability gradient: dimensions legible on the textual surface are annotated consistently, while those requiring inference of intent are contested, delimiting the claims such annotation supports. Scaling the validated extractor to the full Twenty-Four Histories (65,380 comparisons, 5,766 pairs) recovers corpus-level structure a similarity score cannot express. The interpretive composition of citation shows no systematic change across eighteen centuries, yet the same passage is quoted less and less literally. Stability in the aggregate with drift in the individual case is what a cultural-attraction account expects. We release the extraction protocol and the expert-adjudicated benchmark.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!