2605.29940v1 May 28, 2026 cs.AI

피드백을 통한 스트리밍 경험으로부터 LLM이 합성 능력을 학습하도록 하는 방법

Make LLM Learn to Synthesize from Streaming Experiences through Feedback

Zhen Bi
Zhen Bi
Citations: 15
h-index: 2
Jungang Lou
Jungang Lou
Citations: 44
h-index: 3
Zhixuan Chu
Zhixuan Chu
Citations: 3
h-index: 1
Zihao Xue
Zihao Xue
Citations: 6
h-index: 2
Longtao Huang
Longtao Huang
Citations: 31
h-index: 2
Zeyu Yang
Zeyu Yang
Citations: 3
h-index: 1
Bin Zhu
Bin Zhu
Citations: 132
h-index: 5
Yan Wang
Yan Wang
Citations: 12
h-index: 2
Zhen-Hua Hu
Zhen-Hua Hu
Citations: 12
h-index: 2
Xiongtao Zhang
Xiongtao Zhang
Citations: 264
h-index: 9

대규모 언어 모델(LLM)은 합성 데이터 생성에 널리 활용되어, 주석 비용을 크게 절감합니다. 그러나 대부분의 기존 연구는 합성을 독립적인 작업 세트로 취급하며, 더 근본적인 질문인 '모델이 과거 작업에서 얻은 경험을 축적하고 이를 향후 작업에 적용하여 합성 능력을 학습할 수 있는가?'라는 점을 간과합니다. 본 논문에서는 합성 작업이 순차적으로 제공되고 이전 작업의 경험이 향후 합성에 유용한 정보를 제공하는 새로운 환경인 StreamSynth를 소개합니다. 이 환경에 대응하기 위해, SynLearner라는 일반적인 프레임워크를 제안합니다. SynLearner는 합성 모델이 작업 스트림을 통해 재사용 가능한 합성 경험을 습득할 수 있도록 합니다. SynLearner는 각 작업별로 독립적으로 데이터를 생성하는 대신, 모델이 다양한 합성 패턴을 탐색하고, 피드백으로부터 학습하며, 작업의 변화에 따라 샘플 품질과 전체 데이터셋의 다양성을 균형 있게 유지하도록 장려합니다. 여러 벤치마크를 대상으로 실시한 광범위한 실험 결과, SynLearner는 이전 작업에서 얻은 경험을 효과적으로 활용하여 후속 작업에서의 합성 성능을 향상시키며, 일관된 교차 작업 전이성을 보임을 확인했습니다. 이러한 결과는 StreamSynth의 실현 가능성을 입증하고, 합성 데이터 생성 과정을 경험 기반 프로세스로 보고, 작업 스트림으로부터 이점을 얻을 수 있음을 강조합니다.

Original Abstract

Large language models (LLMs) have been widely adopted for synthetic data generation, significantly reducing annotation costs. However, most existing studies treat synthesis as a set of isolated tasks and overlook a more fundamental question: whether a model can learn to synthesize by accumulating experience from past tasks and transferring it to future ones. In this work, we introduce StreamSynth, a new setting in which synthesis tasks arrive sequentially and experience from historical tasks provides informative signals for future synthesis. To address this setting, we propose SynLearner, a general framework that enables synthesis models to acquire reusable synthesis experience over a task stream. Instead of generating data independently for each task, SynLearner encourages the model to explore diverse synthesis patterns, learn from feedback, and balance sample quality with set-level diversity as tasks evolve. Extensive experiments across multiple benchmarks show that SynLearner effectively leverages experience from earlier tasks to improve synthesis performance on later ones, exhibiting consistent cross-task transferability. These findings provide evidence for the feasibility of StreamSynth and highlight synthetic data generation as an experience-driven process that can benefit from task streams.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!