2604.04708v1 Apr 06, 2026 cs.CL

BiST: 문장 구조 및 시제 분류를 위한 표준 벵골어-영어 이중 언어 말뭉치: 어노테이터 간 일치도 분석

BiST: A Gold Standard Bangla-English Bilingual Corpus for Sentence Structure and Tense Classification with Inter-Annotator Agreement

A. Shafi
A. Shafi
Citations: 45
h-index: 4
Abdul Muntakim
Abdul Muntakim
Citations: 36
h-index: 5
Swapnil Kundu Argha
Swapnil Kundu Argha
Citations: 0
h-index: 0
M. A. Moyeen
M. A. Moyeen
Citations: 5
h-index: 1
Shoumik Barman Polok
Shoumik Barman Polok
Citations: 0
h-index: 0

고품질의 이중 언어 자원은, 특히 벵골어와 같은 자원 부족 환경에서 다국어 자연어 처리 기술 발전에 있어 중요한 제약 요인입니다. 이러한 격차를 해소하기 위해, 본 연구에서는 문장 수준의 문법 분류를 위한 엄격하게 관리된 벵골어-영어 말뭉치인 BiST를 소개합니다. BiST는 구문 구조 (단순, 복합, 복합-복합) 및 시제 (현재, 과거, 미래)의 두 가지 기본적인 측면에서 어노테이션이 수행되었습니다. 이 말뭉치는 공개 라이선스를 가진 백과사전 자료와 자연스러운 대화 텍스트를 기반으로 구축되었으며, 체계적인 전처리 및 자동 언어 식별 과정을 거쳐 총 30,534개의 문장, 즉 17,465개의 영어 문장과 13,069개의 벵골어 문장을 포함합니다. 다단계 프레임워크를 통해 세 명의 독립적인 어노테이터가 참여하고, 구조 및 시제 어노테이션에 대해 각각 0.82와 0.88의 Fleiss Kappa ($κ$) 일치도를 사용하여 어노테이션 품질을 보장했습니다. 통계 분석 결과, 실제적인 구문 및 시제 분포가 확인되었으며, 초기 실험 결과는 상호 보완적인 언어별 표현을 활용하는 듀얼 인코더 구조가 강력한 다국어 인코더보다 우수한 성능을 보였습니다. BiST는 벤치마킹 도구로서의 역할뿐만 아니라, 제어된 텍스트 생성, 자동 피드백 생성, 교차 언어 표현 학습을 포함한 문법 모델링 작업을 지원하는 명시적인 언어적 감독 신호를 제공합니다. 이 말뭉치는 이중 언어 문법 모델링을 위한 통일된 자원을 제공하며, 언어학적 기반의 다국어 연구를 촉진합니다.

Original Abstract

High-quality bilingual resources remain a critical bottleneck for advancing multilingual NLP in low-resource settings, particularly for Bangla. To mitigate this gap, we introduce BiST, a rigorously curated Bangla-English corpus for sentence-level grammatical classification, annotated across two fundamental dimensions: syntactic structure (Simple, Complex, Compound, Complex-Compound) and tense (Present, Past, Future). The corpus is compiled from open-licensed encyclopedic sources and naturally composed conversational text, followed by systematic preprocessing and automated language identification, resulting in 30,534 sentences, including 17,465 English and 13,069 Bangla instances. Annotation quality is ensured through a multi-stage framework with three independent annotators and dimension-wise Fleiss Kappa ($κ$) agreement, yielding reliable and reproducible labels with $κ$ values of 0.82 and 0.88 for structural and temporal annotation, respectively. Statistical analyses demonstrate realistic structural and temporal distributions, while baseline evaluations show that dual-encoder architectures leveraging complementary language-specific representations consistently outperform strong multilingual encoders. Beyond benchmarking, BiST provides explicit linguistic supervision that supports grammatical modeling tasks, including controlled text generation, automated feedback generation, and cross-lingual representation learning. The corpus establishes a unified resource for bilingual grammatical modeling and facilitates linguistically grounded multilingual research.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!