논쟁을 보상으로: 강화학습 후처리(Post-Training)를 이용한 과학적 아이디어 창출을 위한 다중 에이전트 보상 시스템
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
대규모 언어 모델(LLM)은 과학적 아이디어 창출을 자동화하는 데 잠재력을 보여주었지만, 반복적인 프롬프팅이나 복잡한 다중 에이전트 아키텍처에 의존하는 현재 방법은 종종 환각 현상이나 계산 효율성 저하 문제를 겪습니다. 강화학습(RL)을 이 개방형 영역에 적용하는 데 있어 중요한 병목 현상은 '보상 해킹(reward hacking)'입니다. 이는 모델이 불완전한 평가 지표를 이용하여 점수를 극대화하지만, 진정한 과학적 혁신을 산출하지 않는 현상을 말합니다. 이러한 한계점을 해결하기 위해, 우리는 고품질의 과학적 아이디어 생성에 특화된 RL 프레임워크를 제안합니다. 우리는 방법론적 검증과 구현 세부 사항을 분리하면서, 보상 해킹에 강한 엄격한 이진 보상을 제공하는, 최초의 다중 에이전트 보상 함수를 설계했습니다. 이 희소한 신호에 효과적으로 최적화하기 위해, 인공적인 길이 편향을 완화하기 위해 Group Relative Policy Optimization의 편향되지 않은 변형을 사용합니다. 우리는 IC LR-320이라는, IC LR 2024 발표 내용에서 추출한 문제-해결 쌍 데이터 세트를 기반으로 학습을 진행했습니다. 실험 결과, 우리의 프레임워크가 전문가가 평가한 참신성, 실현 가능성, 효과성 측면에서 최첨단 모델보다 훨씬 뛰어난 성능을 보이는 것을 확인했습니다.
Large Language Models (LLMs) have demonstrated potential in automating scientific ideation, yet current approaches relying on iterative prompting or complex multi-agent architectures often suffer from hallucination or computational inefficiency. A critical bottleneck in applying Reinforcement Learning (RL) to this open-ended domain is reward hacking -- where models exploit imperfect evaluation proxies to maximize scores without producing genuine scientific innovation. To address these limitations, we propose an RL framework explicitly tailored for high-quality scientific idea generation. We propose the first multi-agent reward function designed to serve as a judge, decoupling methodological validation from implementation details while providing strict binary rewards that are robust to reward hacking. To effectively optimize against this sparse signal, we utilize an unbiased variant of Group Relative Policy Optimization to mitigate artificial length bias. We grounded our training in ICLR-320, a curated dataset of problem-solution pairs extracted from ICLR 2024 proceedings. Experiments demonstrate that our framework significantly outperforms state-of-the-art baselines across expert-evaluated metrics of novelty, feasibility, and effectiveness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.