2605.27858v1 May 27, 2026 cs.CL

DecomposeRL: 반지도 학습 및 추적 가능한 주장에 대한 검증을 위한 유용하고, 정보력이 있으며, 다양한 질문 학습

DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervised, Traceable Claim Verification

Shubhashis Roy Dipta
Shubhashis Roy Dipta
University of Maryland, Baltimore County
Citations: 130
h-index: 6
Ankur Padia
Ankur Padia
Citations: 338
h-index: 8
Francis Ferraro
Francis Ferraro
Citations: 21
h-index: 1

주장 검증은 정확하지만 검사할 수 없는 추적 정보를 제공하는 통합 분류기와, 검사 가능한 추적 정보를 생성하지만 벤치마크 데이터 세트에서 성능이 떨어지는 분해 기반 방법으로 나뉩니다. 우리는 정확하고 검사 가능한 추적 정보를 제공하는 주장 검증 시스템인 DecomposeRL을 제안합니다. DecomposeRL은 분해를 GRPO로 학습된 강화 학습 정책으로 정의하며, 다면적인 보상 집합을 사용하여 완전 지도 및 반지도 학습을 모두 지원합니다. DecomposeRL은 GRPO의 높은 훈련 비용 문제를 해결하기 위해 데이터 선별 과정을 통해 115,000개의 사실 검증 주장을 5,000개의 학습 신호가 풍부한 부분집합으로 축소했습니다. 실험 결과, 약 5,000개의 선별된 주장 데이터를 사용하여 완전 지도 방식으로 훈련된 DecomposeRL-7B 정책은 생물 의학, 정치, 과학 및 일반 도메인 주장을 포함하는 11개의 주장 검증 벤치마크에서 86.3%의 인접 영역 정확도와 69.8%의 외부 영역 균형 정확도를 달성했습니다. DecomposeRL 모델은 크기가 4배 작음에도 불구하고 320억 개의 기본 모델 및 GPT-4.1-mini와 동등한 성능을 보이며, 전체 데이터 중 10%만 레이블이 지정된 반지도 학습 환경에서는 기본 모델보다 더 뛰어난 성능을 보였습니다. 코드, 데이터 및 모델은 https://dipta007.github.io/DecomposeRL에서 확인할 수 있습니다.

Original Abstract

Claim verification splits between end-to-end classifiers that are accurate but yields no inspectable traces, and decomposition-based methods produce inspectable traces but lag performance on benchmark datasets. We propose DecomposeRL an accurate claim-verifier that produce inspectable traces. DecomposeRL frames decomposition as an RL policy trained with GRPO and a multi-faceted reward ensemble, enabling both fully supervised and semi-supervised learning from unlabeled claims. DecomposeRL addresses the prohibitive training cost of GRPO with a data-curation funnel that distills 115K fact-verification claims into a compact, learning-signal-dense subset of 5K claims. We show that a DecomposeRL-7B policy trained with full supervision on only ~5K curated claims achieves 86.3 in-domain and 69.8 out-of-domain balanced accuracy across 11 claim-verification benchmarks containing biomedical, political, scientific, and general-domain claims. Despite being 4x smaller, it matches 32B baselines and GPT-4.1-mini, and it further outperforms baselines in a semi-supervised setting with only 10% labeled claims data. Code, data, and models are available at https://dipta007.github.io/DecomposeRL

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!