2608.01862v1 Aug 03, 2026 cs.AI

Wnuan: 질문 답변을 위한 단계별 사후 학습 - 독점적인 기업 지식 기반

Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

Xiaofeng Shi
Xiaofeng Shi
Citations: 51
h-index: 4
Xiaosong Qiu
Xiaosong Qiu
Citations: 8
h-index: 2
Wenxin Ma
Wenxin Ma
Citations: 0
h-index: 0
Qian Kou
Qian Kou
Citations: 20
h-index: 3
Yiming Pan
Yiming Pan
Citations: 22
h-index: 3
Longbin Yu
Longbin Yu
Citations: 16
h-index: 2
Ying Liu
Ying Liu
Citations: 0
h-index: 0
Haiping Wang
Haiping Wang
Citations: 0
h-index: 0
Hua Zhou
Hua Zhou
Citations: 1,199
h-index: 5

기업용 질의응답 시스템은 모델이 일반적인 능력을 희생하지 않고도 독점적인 지식을 습득하도록 요구합니다. 본 논문에서는 Wnuan이라는 세 단계를 거치는 파이프라인을 소개합니다. 이 파이프라인은 문서로부터 작업 중심의 감독 데이터를 구축하고, 일반 데이터 재학습을 통해 지도 학습을 수행하며, 잔여 오류에 대해 강화 학습을 적용합니다. 707개의 질문으로 구성된 WnuanBench 데이터셋에서, 초기 모델(32B)의 정답률이 사전 적응 시 52.76%였으나, 지도 미세 조정(SFT) 후에는 80.06%, 강화 학습 적용 후에는 91.51%로 향상되었습니다. 동일한 100회 업데이트 환경에서, 잔여 오류 샘플링은 전체 풀 샘플링 및 크기 일치 랜덤 샘플링보다 각각 3.11점과 2.97점을 높였습니다. 소스 클러스터 부트스트랩 구간은 두 방법 모두 양수로 유지되었으며, 동일 도메인 검증 데이터셋에서도 순서가 유지되었습니다. 일반적인 벤치마크 성능은 전체적으로 5.17점이 감소했는데, 이는 지시 사항 준수에 집중된 결과입니다. 자동 평가 시스템은 계층화된 Wnuan-Inst 응답 샘플의 90.5%에 대해 해당 분야 전문가의 판단과 일치했습니다. 이러한 결과는 단계별 기업 환경 적응을 통해 얻을 수 있는 이점과 일반적인 능력 저하라는 비용을 동시에 보여줍니다.

Original Abstract

Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. Under a matched 100-update protocol, residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points, respectively. Source-cluster bootstrap intervals remain above zero for both contrasts, and a same-domain validation set preserves the ordering. The general-benchmark average decreases by 5.17 points across the route, concentrated in instruction following. The automatic evaluation ensemble agrees with an authoritative domain expert on 90.5% of a stratified Wnuan-Inst response sample. These results characterize both the gains and the general-capability cost of staged enterprise adaptation.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!