다중 영역 저자원 광학 문자 인식(OCR)을 위한 다중 전문가 라우팅: 만주어 사례 연구
Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study
본 논문에서는 제한된 레이블 데이터에도 불구하고 정규체, 흘림체, 그리고 궁궐 문서를 작성하는 데 사용된 준서예 필기체를 포함하여 다양한 시각적으로 구별되는 서체 스타일을 수용해야 하는 역사적 만주어 OCR 문제를 다룬다. 우리는 반복적인 미세 조정 과정을 통해 얻은 체크포인트를 도메인 전문가로 재사용하고, 페이지 수준의 경량 이미지 분류기를 사용하여 페이지를 시각적 스타일에 따라 분배하는 다중 전문가 시스템을 연구한다. 체크포인트 풀에 적합한 전문 지식이 없을 경우, 해당 도메인을 위한 추가적인 전문가를 학습시킨다. 세 개의 고정된 테스트 데이터셋에서, 라우팅된 시스템은 각 스타일별로 선택된 전문가와 일치하며, 정규체는 0.30%, 궁궐 문서는 1.57%, 흘림체는 4.83%의 CER(Character Error Rate)을 보인다. 라우터는 페이지 수준에서 99.3%의 도메인 정확도를 달성했으며, 도메인 레이블 오라클과 동일한 정밀도로 일치한다. 세 개의 선택된 전문가 중 두 명은 최종 도메인을 위해 특별히 학습되지 않았으며, 흘림체 전문가는 해당 도메인을 목표로 학습되었다. 본 논문에서는 비교 가능성을 확보하기 위해 평가 프로토콜, 라우터 설계 및 페이지별 예측 결과를 상세하게 제시한다.
Historical Manchu OCR must accommodate various visually distinct writing styles, including regular script, running script, and the semi-cursive chancery hand used in palace memorials, despite limited labeled data. We study a multi-expert system that reuses checkpoints from an iterative fine-tuning process as domain specialists and uses a lightweight page-level image classifier to dispatch pages by visual style. When the checkpoint pool lacks a suitable specialist, we train an additional expert for that domain. On three frozen test sets, the routed system matches the selected specialist for each style at two-decimal precision: 0.30 percent CER on regular script, 1.57 percent on memorials, and 4.83 percent on running script. The router achieves 99.3 percent page-level domain accuracy and matches the domain-label oracle at the same precision. Two of the three selected specialists were not trained specifically for their final domain; only the running-script expert was trained with that domain as its target. We report the evaluation protocol, router design, and per-page predictions to make the comparison reproducible.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.