Wnuan: 질문 답변을 위한 단계별 사후 학습 - 독점적인 기업 지식 기반
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
기업용 질의응답 시스템은 모델이 일반적인 능력을 희생하지 않고도 독점적인 지식을 습득하도록 요구합니다. 본 논문에서는 Wnuan이라는 세 단계를 거치는 파이프라인을 소개합니다. 이 파이프라인은 문서로부터 작업 중심의 감독 데이터를 구축하고, 일반 데이터 재학습을 통해 지도 학습을 수행하며, 잔여 오류에 대해 강화 학습을 적용합니다. 707개의 질문으로 구성된 WnuanBench 데이터셋에서, 초기 모델(32B)의 정답률이 사전 적응 시 52.76%였으나, 지도 미세 조정(SFT) 후에는 80.06%, 강화 학습 적용 후에는 91.51%로 향상되었습니다. 동일한 100회 업데이트 환경에서, 잔여 오류 샘플링은 전체 풀 샘플링 및 크기 일치 랜덤 샘플링보다 각각 3.11점과 2.97점을 높였습니다. 소스 클러스터 부트스트랩 구간은 두 방법 모두 양수로 유지되었으며, 동일 도메인 검증 데이터셋에서도 순서가 유지되었습니다. 일반적인 벤치마크 성능은 전체적으로 5.17점이 감소했는데, 이는 지시 사항 준수에 집중된 결과입니다. 자동 평가 시스템은 계층화된 Wnuan-Inst 응답 샘플의 90.5%에 대해 해당 분야 전문가의 판단과 일치했습니다. 이러한 결과는 단계별 기업 환경 적응을 통해 얻을 수 있는 이점과 일반적인 능력 저하라는 비용을 동시에 보여줍니다.
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. Under a matched 100-update protocol, residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points, respectively. Source-cluster bootstrap intervals remain above zero for both contrasts, and a same-domain validation set preserves the ordering. The general-benchmark average decreases by 5.17 points across the route, concentrated in instruction following. The automatic evaluation ensemble agrees with an authoritative domain expert on 90.5% of a stratified Wnuan-Inst response sample. These results characterize both the gains and the general-capability cost of staged enterprise adaptation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.