2606.24172v1 Jun 23, 2026 cs.CL

인도어 자연어 처리의 파니니 문법 기반 접근

A Pāninian Foundation for Indic Language Processing

L. Varshney
L. Varshney
Citations: 0
h-index: 0
Ritwik Banerjee
Ritwik Banerjee
Citations: 0
h-index: 0

10억 명이 넘는 사람들이 인도어를 사용하고 있지만, 이들을 위한 자연어 처리 인프라는 여전히 단편적이고 발전이 미흡한 상태입니다. 이는 구조적인 문제입니다. 해당 분야에서는 도구와 벤치마크를 개별 언어나 작은 어족 그룹에 맞춰 개발하며, 각 언어마다 별도의 분석기, 파서 및 데이터셋을 구축하고 다음 언어로 넘어갈 때마다 처음부터 다시 시작합니다. 이러한 접근 방식은 깊은 수준의 규칙성을 간과합니다. 수천 년 동안 산스크리트어를 중심으로 인도어들은 진화해 왔으며, 이는 파니니의 문법인 아스타디아이에 공식화된 형태론-구문 구조를 공유하게 되었습니다. 이 공통 구조는 어족 경계를 넘어 언어들을 하나로 묶습니다. 우리는 이러한 파니니 문법 기반 프레임워크가 해당 분야에 부족했던 통합적인 계산 모델을 제공하며, 이를 명시적으로 반영한 벤치마크는 인도어 시스템의 정확성, 데이터 효율성 및 적용 가능성을 향상시켜, 겉보기에 서로 다른 희소한 인도어 자원들을 하나의 풍부한 메타 언어 기반으로 효과적으로 통합할 수 있을 것이라고 주장합니다. 우리는 이러한 공유된 구조를 명확하게 드러내고 측정 가능하도록 만들기 위해 네 부분으로 구성된 벤치마크 세트를 제안합니다. 또한, 본 연구는 해석 가능성 연구에 중요한 질문을 던집니다. 즉, 이 언어들로 학습된 신경망 모델이 자체적으로 파니니의 범주를 표현하게 될 수 있는지 여부에 대한 문제입니다.

Original Abstract

More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragmented and underdeveloped. The cause is structural: the field organizes its tools and benchmarks around individual languages or small subsets of genealogical language families, building separate analyzers, parsers, and datasets for each language and starting over for the next. This overlooks a deep regularity. Through more than two millennia of convergence around Sanskrit, Indic languages came to share a morphosyntactic architecture formalized in Pānini's grammar, the Astādhyāyī. This cuts across genealogical lines, uniting languages through a common framework. We argue that this Pāninian framework supplies a unifying computational architecture the field has lacked, and that benchmarks grounded explicitly in it would make Indic language systems more accurate, more data-efficient, and more transferable, effectively merging many apparently disparate and sparse Indic language resources into a single high-resource metalanguage bedrock. We propose a four-part benchmark suite to render this shared architecture explicit, measurable, and ready to be leveraged for practical applications. Moreover, we underscore the question it raises for interpretability research: whether neural models trained on these languages come to represent Pānini's categories on their own.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!