2607.01846v1 Jul 02, 2026 cs.AI

CLAP: 도메인 에이전트 사후 학습을 위한 폐쇄 루프 기반 훈련, 평가 및 배포 제어

CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training

Feng Tian
Feng Tian
Citations: 16,223
h-index: 6
Fangfei Li
Fangfei Li
Citations: 0
h-index: 0
Chenyang Zhao
Chenyang Zhao
Citations: 0
h-index: 0
Zhiyue Zheng
Zhiyue Zheng
Citations: 0
h-index: 0
Long Wang
Long Wang
Citations: 0
h-index: 0
Lv Guo
Lv Guo
Citations: 0
h-index: 0

도메인 에이전트는 종종 노이즈가 많은 비즈니스 데이터, 불확실한 사후 학습 효과, 오프라인/애플리케이션 불일치, 그리고 어댑터 배포 위험에 직면합니다. 본 논문에서는 CLAP (Closed-Loop Agent Post-training)이라는 폐쇄 루프 방식을 제시하며, 이 방식은 비즈니스 데이터를 구조화된 SFT 샘플, 의사 결정 선호도 샘플, 검증 데이터 세트, 위험 진단 및 배포 게이트 기록으로 변환합니다. CLAP은 데이터 유효성 검사, 목표/증거 정규화, 보상/KL 진단, 오프라인 게이트, 그리고 애플리케이션 체인 리플레이를 결합하여 어댑터가 대상 애플리케이션 체인에 적합한지 판단합니다. 다섯 개의 익명화된 제조 시나리오 데이터 세트에서 QLoRA 스타일의 LoRA-SFT는 미미한 평균적인 성능 향상을 보여주었습니다: 전체 점수는 0.0098, 합격률은 0.0240, 증거 정확도는 0.0280이 증가했으며, 환각 현상과 잘못된 사실은 감소했습니다. 그러나 5개의 데이터 세트 중 단 3개만 성능이 향상되었고, 일부 데이터 세트는 오히려 성능이 저하되었으며, GRPO는 높은 KL 위험을 드러냈습니다. 또한 애플리케이션 체인 리플레이 실험 결과, 사실 추출에는 RAG가 필수적임을 보여주었습니다. 동일한 3B 모델과 100개의 리플레이 케이스를 사용했을 때, 애플리케이션-RAG에 특화된 LoRA-SFT 어댑터는 기본 모델 + RAG보다 가치, 핵심 필드 및 답변-증거 문서/페이지 매칭 측면에서 성능이 향상되었지만, 지연 시간이 증가했습니다. 이러한 결과는 도메인 에이전트의 사후 학습을 훈련 완료 여부나 단일 오프라인 점수에 의존하는 것이 아니라 통합된 데이터-훈련-평가-배포 루프를 통해 관리해야 함을 뒷받침합니다.

Original Abstract

Domain agents often face noisy business data, uncertain post-training gains, offline/application mismatch, and adapter-release risk. This paper presents CLAP (Closed-Loop Agent Post-training), a closed-loop method that converts business data into structured SFT samples, decision-preference samples, holdout sets, risk diagnostics, and release-gate records. CLAP combines data validation, target/evidence normalization, reward/KL diagnosis, offline gates, and application-chain replay to decide whether an adapter is suitable for the target application chain. On five anonymized manufacturing-scenario batches, QLoRA-style LoRA-SFT yields modest average gains: overall score increases by 0.0098, pass rate by 0.0240, and evidence accuracy by 0.0280, while hallucination and wrong facts decrease. Yet only 3 of 5 batches improve, some batches regress, and GRPO exposes high KL risks. Application-chain replay further shows that RAG is necessary for factual extraction; under the same 3B backbone and 100 replay cases, an application-RAG-oriented LoRA-SFT adapter improves value, core fields, and answer-evidence doc/page matching over base+RAG, but increases latency. These results support managing domain-agent post-training through an integrated data-training-evaluation-release loop rather than relying on training completion or a single offline score.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!