TongGuOCR: 레이아웃 인식 및 토큰 증강 기반의 중국 역사 문서 광학 문자 인식 멀티모달 대규모 언어 모델
TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents
중국 역사 문서는 귀중한 문화 유산을 담고 있지만, 많은 자료가 스캔된 이미지 형태로만 존재하여 전체 텍스트 검색, 비교 분석 및 계산적 분석이 어렵습니다. 광학 문자 인식(OCR) 기술은 이러한 문제를 해결할 수 있지만, 복잡한 레이아웃, 희귀 문자, 그리고 다양한 읽기 순서 때문에 정확한 전사 작업은 여전히 어려운 과제입니다. 본 논문에서는 중국 역사 문서의 OCR을 위한 레이아웃 인지 및 토큰 증강 기반의 멀티모달 대규모 언어 모델(MLLM)인 TongGuOCR을 제안합니다. 먼저, 레이아웃 인식 전처리 모듈은 로컬 문맥을 유지하면서 영역 간 간섭을 줄이기 위해 일관성 있는 인식 블록을 구성하고 개선합니다. 둘째, 토큰 증강 인식 모듈은 두 가지 상호 보완적인 수준에서 전사 대상을 확장합니다. 문자 레벨의 어휘 확장은 각 희귀 글자에 직접 하나의 토큰 표현을 제공하여 디코딩 경로를 단축시키고, 라인-투-라인 전환 모델링은 이산적인 공간 이동 토큰을 주입하여 정확한 좌표 없이 복잡한 읽기 경로를 따라 디코더를 안내합니다. 두 개의 중국 역사 문서 OCR 벤치마크에서 수행된 실험 결과, TongGuOCR은 대표적인 기존의 특정 작업 전용 OCR 모델, 범용 MLLM 및 OCR 특화 MLLM보다 우수한 성능을 보였습니다. 특히 더 어려운 M5HisDoc 벤치마크에서 TongGuOCR은 AR 점수를 93.76으로 달성했으며, NED를 10.43에서 6.15로, RO-ED를 7.53에서 3.49로 각각 감소시켜 각 지표별 최고 경쟁 성능을 능가했습니다. 온라인 데모는 https://jzzh2004.github.io/TongGuOCR 에서 확인할 수 있습니다.
Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bridge this gap, but accurate transcription remains challenging because historical documents often contain complex layouts, rare characters, and nontrivial reading orders. We propose TongGuOCR, a layout-aware and token-augmented multimodal large language model (MLLM) for OCR of Chinese historical documents. First, a Layout-Aware Preprocessing module constructs and refines locally coherent recognition blocks to preserve local context while reducing interference across regions. Second, a Token-Augmented Recognition module augments the transcription target at two complementary levels: character-level vocabulary expansion gives each rare glyph a direct one-token representation and shortens its decoding path, while line-to-line transition modeling injects discrete spatial displacement tokens that guide the decoder along complex reading paths without requiring precise coordinates. Experiments on two Chinese historical document OCR benchmarks show that TongGuOCR outperforms representative traditional task-specific OCR models, general-purpose MLLMs, and OCR-oriented MLLMs. On the more challenging M5HisDoc benchmark, TongGuOCR achieves 93.76 AR and reduces NED from 10.43 to 6.15 and RO-ED from 7.53 to 3.49 relative to the best competing score for each metric. An online demo is available at https://jzzh2004.github.io/TongGuOCR.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.