2603.21728v1 Mar 23, 2026 cs.AI

EvoIdeator: 체크리스트 기반 강화 학습을 통한 과학적 아이디어의 진화

EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning

Lu Zhou
Lu Zhou
Citations: 109
h-index: 5
Andreas Sauter
Andreas Sauter
Citations: 21
h-index: 3
Yuyue Zhao
Yuyue Zhao
Citations: 23
h-index: 2
Jacopo Urbani
Jacopo Urbani
Citations: 2,002
h-index: 24
Wenxiang Hu
Wenxiang Hu
Citations: 63
h-index: 3
Zaiqiao Meng
Zaiqiao Meng
Citations: 123
h-index: 6
Xiaohu Yan
Xiaohu Yan
Citations: 22
h-index: 2
Yougang Lyu
Yougang Lyu
Citations: 271
h-index: 10

과학적 아이디어 생성은 자율적인 지식 발견의 핵심이지만, 초기 개념을 고품질 연구 제안으로 발전시키기 위한 반복적인 진화 과정은 대규모 언어 모델(LLM)에게 여전히 큰 과제입니다. 기존의 강화 학습(RL) 패러다임은 종종 전반적인 품질 점수를 제공하지만 실행 가능한 세부 정보를 제공하지 못하는 rubic 기반의 스칼라 보상에 의존합니다. 반대로, 언어 기반의 개선 방법은 일반적으로 추론 시간에 프롬프트를 사용하는 방식으로 제한되며, 그러한 비판을 내재화하도록 명시적으로 최적화되지 않은 모델을 대상으로 합니다. 이러한 격차를 해소하기 위해, 우리는 체크리스트 기반의 피드백을 활용하여 과학적 아이디어의 진화를 촉진하는 프레임워크인 **EvoIdeator**를 제안합니다. EvoIdeator는 구조화된 평가 모델을 활용하여 두 가지 상호 보완적인 신호를 생성합니다. (1) 다차원 최적화를 위한 extit{어휘적 보상} 및 (2) 근거, 실현 가능성 및 방법론적 엄격성에 대한 스팬 레벨의 비판을 제공하는 extit{세분화된 언어 피드백}. 이러한 신호를 RL 루프에 통합함으로써, 정책이 최적화 및 추론 과정 모두에서 체계적으로 정확한 피드백을 활용하도록 훈련합니다. 광범위한 실험 결과, Qwen3-4B를 기반으로 구축된 EvoIdeator가 주요 과학적 지표에서 훨씬 더 큰 최첨단 모델보다 뛰어난 성능을 발휘하는 것으로 나타났습니다. 더욱 중요하게는, 학습된 정책은 추가적인 미세 조정 없이 다양한 외부 피드백 소스에 대한 강력한 일반화 능력을 보여주며, 자체 개선이 가능한 자율적인 아이디어 생성을 위한 확장 가능하고 엄격한 경로를 제공합니다.

Original Abstract

Scientific idea generation is a cornerstone of autonomous knowledge discovery, yet the iterative evolution required to transform initial concepts into high-quality research proposals remains a formidable challenge for Large Language Models (LLMs). Existing Reinforcement Learning (RL) paradigms often rely on rubric-based scalar rewards that provide global quality scores but lack actionable granularity. Conversely, language-based refinement methods are typically confined to inference-time prompting, targeting models that are not explicitly optimized to internalize such critiques. To bridge this gap, we propose \textbf{EvoIdeator}, a framework that facilitates the evolution of scientific ideas by aligning the RL training objective with \textbf{checklist-grounded feedback}. EvoIdeator leverages a structured judge model to generate two synergistic signals: (1) \emph{lexicographic rewards} for multi-dimensional optimization, and (2) \emph{fine-grained language feedback} that offers span-level critiques regarding grounding, feasibility, and methodological rigor. By integrating these signals into the RL loop, we condition the policy to systematically utilize precise feedback during both optimization and inference. Extensive experiments demonstrate that EvoIdeator, built on Qwen3-4B, significantly outperforms much larger frontier models across key scientific metrics. Crucially, the learned policy exhibits strong generalization to diverse external feedback sources without further fine-tuning, offering a scalable and rigorous path toward self-refining autonomous ideation.

1 Citations
0 Influential
12 Altmetric
61.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!