합성 후처리 데이터 정제 과정에서 출처 정보를 기반으로 한 필터링 및 적응적 복구
Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation
인공적으로 생성된 데이터를 활용하는 후처리 파이프라인은 일반적으로 보상 모델 또는 LLM 평가기를 사용하여 생성된 샘플을 필터링하지만, 두 가지 요소는 종종 함께 고려되지 않습니다. 첫째, 필터링 신호가 각 생성 과정을 유발한 원본 데이터에 기반하고 있는지 여부이고, 둘째, 거부된 샘플을 영구적으로 폐기하는 대신 체계적으로 복구할 수 있는지 여부입니다. 본 연구에서는 적대적으로 삽입된 데이터를 사용하여 진실성 실패 라벨을 제공하고, 다양한 필터링 설정, 복구 전략 및 생성 모델 규모에 따른 두 가지 질문에 대한 통제된 실험을 수행했습니다. 그 결과, 정확한 원본 정보는 더 강력한 평가기를 사용할 때 신뢰성을 높이는 데 도움이 된다는 것을 확인했으며, 환각(hallucination)과 보상 기반 필터링은 대부분 서로 다른 샘플 집단을 거부하므로 두 가지 모두 필요하다는 것을 알았습니다. 또한, 실패 진단과 목표 지향적인 재생성 방법을 결합한 적응적 복구 파이프라인은 단순한 리샘플링보다 높은 수율, 복구율 및 삽입 기억률을 달성했습니다. 하위 작업(downstream)의 미세 조정 품질은 주로 생성 모델의 규모에 의해 결정되며, 필터링 및 복구 조건은 중요한 역할을 하지만 상대적으로 덜 중요한 영향을 미칩니다.
Synthetic post-training pipelines commonly filter generated samples with reward models or holistic LLM judges, yet two practices remain rarely examined together: whether the filtering signal is grounded in the source evidence that induced each generation, and whether rejected samples can be systematically recovered rather than permanently discarded. We present a controlled study of both questions across gate configurations, recovery strategies, and generator scales, using adversarially injected corpora to provide ground-truth failure labels. We find that exact source provenance improves faithfulness gating for stronger judges, that hallucination and reward gates reject largely disjoint sample populations making both necessary, and that an adaptive recovery pipeline combining failure diagnosis with targeted regeneration achieves higher yield, recovery rate, and injection recall than naive resampling. Downstream fine-tuning quality is driven primarily by generator scale, with filtration and recovery conditions contributing meaningfully but secondarily.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.