국가 표준 문서의 규칙 기반 검토를 위한 LLM 성능 평가 및 개선
Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents
대규모 언어 모델(LLM)은 점점 더 복잡한 전문 업무를 지원하고 있지만, 규칙 기반 문서를 검토하는 능력에 대한 평가는 아직 부족합니다. 중국 GB/T 표준과 같은 국가 표준 문서는 대표적인 테스트 환경을 제공합니다. 이 문서들은 분량이 길고, 구조화되어 있으며, 범위, 용어, 규범적 표현 및 일관성 등에 대한 명확한 규칙으로 통제됩니다. 기존의 벤치마크는 주로 도메인 지식과 질문 응답에 초점을 맞추고 있어, 전문 문서의 고유한 품질 검토 측면은 간과되는 경향이 있습니다. 이러한 검토는 일반적으로 인간 전문가에 크게 의존하며, 이는 비용이 많이 들고 확장하기 어렵습니다. 이러한 격차를 해소하기 위해, 본 연구에서는 국가 표준 문서의 체계적인 검토를 위한 최초의 벤치마크인 GB/T-Bench를 소개합니다. GB/T Review Taxonomy는 문서 구조, 범위 일관성, 규범적 모드, 용어 일관성 및 규범적 참조를 포괄하는 계층적 스키마이며, 25가지 진단 가능한 오류 유형을 포함합니다. 제어 가능한 반례 생성 메커니즘은 결정적인 규칙과 제약 조건이 있는 LLM 재작성을 결합하여 488개의 문서를 처리하고, 평가를 위한 7,306개의 추적 가능한 검토 오류 사례를 생성합니다. 또한, 오류 위치, 검토 차원 및 오류 유형에 대한 정확한 일치를 요구하는 진단 중심의 평가 프로토콜을 개발했습니다. 더 나아가, GB/T-Reviewer라는 다중 에이전트 프레임워크를 제안합니다. 이 프레임워크는 검토 지식을 전문 기술로 변환하고 글로벌 검사, 목표 진단, 규칙 스캔 및 결과 확인을 조정합니다. 14개의 주류 LLM에 대한 실험 결과, 인간과 LLM 간의 상당한 격차가 있음을 보여주었습니다. 가장 강력한 모델은 0.3280의 CMCS를 달성하는 데 그친 반면, 전문가의 경우 0.6640을 기록했습니다. GB/T-Reviewer는 CMCS를 0.5094로 향상시켜, 규칙 기반 문서 검토에서 체계적인 기술 조정의 가치를 입증했습니다. 본 연구는 표준화 및 기타 고위험 문서 분야에서 신뢰할 수 있는 AI 개발에 기여할 것입니다.
Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency. Existing benchmarks focus on domain knowledge and question answering, largely overlooking intrinsic quality review for professional documents. Such reviews rely heavily on human experts, making them costly and difficult to scale. To bridge this gap, we introduce GB/T-Bench, the first benchmark for the structured review of national standard documents. Its GB/T Review Taxonomy is a hierarchical schema covering document structure, scope alignment, normative modality, terminology consistency, and normative references, with 25 diagnosable error types. A controllable counterexample generation mechanism combines deterministic rules and constrained LLM rewriting to process 488 documents into 7,306 traceable review error instances for evaluation. We also develop a diagnosis-oriented evaluation protocol requiring exact matches on error location, review dimension, and error type, plus document-level coverage metrics. We further propose GB/T-Reviewer, a multi-agent framework that converts review knowledge into specialized skills and coordinates global inspection, targeted diagnosis, rule scanning, and result verification. Experiments with 14 mainstream LLMs reveal a substantial human-LLM gap: the strongest model achieves only 0.3280 CMCS versus 0.6640 for experts. GB/T-Reviewer raises the best CMCS to 0.5094, showing the value of structured skill coordination for rule-intensive document review. This work paves the way for trustworthy AI in standardization and other high-stakes document domains.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.