2607.00531v1 Jul 01, 2026 cs.LG

Active-GRPO: 분자 최적화를 위한 적응형 모방 및 자기 개선 추론

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization

Le Cong
Le Cong
Citations: 12
h-index: 3
Thomas S. Brettin
Thomas S. Brettin
Citations: 617
h-index: 8
Mingxuan Cao
Mingxuan Cao
Citations: 22
h-index: 3
Xuefeng Liu
Xuefeng Liu
Citations: 92
h-index: 6
Qinan Huang
Qinan Huang
Citations: 6
h-index: 1
Rick L. Stevens
Rick L. Stevens
Citations: 51
h-index: 4

대규모 언어 모델의 중요한 능력 중 하나인 과학적 추론 능력을 향상시키는 것은 여전히 큰 과제입니다. 본 연구에서는 지시 기반 분자 최적화 문제를 중심으로 이 문제에 대해 탐구합니다. 기존의 답변 중심 지도 학습(SFT) 방식은 다단계 추론을 효과적으로 처리하지 못하며, 검증 가능한 보상을 활용한 강화 학습(RLVR)은 희소한 피드백으로 인해 어려움을 겪습니다. 레퍼런스 기반 정책 최적화(Reference-guided Policy Optimization)는 데이터셋에서 제공하는 레퍼런스를 사용하여 정책 업데이트를 안정화하지만, 그 효과는 레퍼런스의 품질에 크게 의존합니다. 레퍼런스가 좋지 않거나 일치하지 않으면 성능 향상에 한계가 있습니다. 이러한 한계를 극복하기 위해, 본 연구에서는 '적극적인 추론(active reasoning)'이라는 새로운 패러다임을 제안합니다. 이 패러다임은 정책이 각 인스턴스별로 레퍼런스를 모방할지 또는 자체적으로 발견한 내용을 강화학습할지를 능동적으로 결정하며, 동시에 지속적으로 모방 대상을 업그레이드합니다. 이러한 패러다임을 'Active Group Relative Policy Optimization (Active-GRPO)'이라는 방식으로 구현했습니다. Active-GRPO는 두 가지 주요 메커니즘, 즉 '적극적인 모방-강화(active imitate-reinforce)'와 '적극적인 레퍼런스 업데이트(active referencing)'로 구성됩니다. '적극적인 모방-강화'는 정책이 자체적으로 생성한 후보 분자가 레퍼런스를 능가할 때까지는 레퍼런스를 모방하는 학습을 수행하고, 그 이후에는 강화 학습을 통해 자체 개선을 진행합니다. '적극적인 레퍼런스 업데이트'는 현재까지 가장 우수한 정책에서 생성된 후보 분자를 찾아 레퍼런스로 대체함으로써, 지속적으로 레퍼런스를 업그레이드하여 학습 과정 동안 레퍼런스가 정보 제공 역할을 계속 수행하도록 합니다 (제한적인 역할이 아닌). TOMG-Bench MOLOPT 데이터셋에 대한 실험 결과, Active-GRPO는 GRPO(0.0959) 및 RePO(0.1665)에 비해 평균 SRxSim 값을 0.1773으로 향상시켰으며, 이는 세 가지 다른 초기값 설정 하에서 통계적으로 유의미한 결과입니다. 또한 LogP, MR, QED 값에서도 상당한 개선을 보였습니다.

Original Abstract

Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such reasoning remains a key open challenge. We study this problem in instruction-based molecular optimization, where answer-only supervised fine-tuning (SFT) collapses multi-step reasoning and reinforcement learning with verifiable rewards (RLVR) suffers from sparse feedback. Reference-guided Policy Optimization mitigates both by anchoring policy updates to dataset-provided references, but its effectiveness is tightly coupled to reference quality: weak or misaligned references impose a performance ceiling. To overcome this ceiling, we propose active reasoning, a paradigm in which the policy actively decides, on a per-instance basis, when to imitate a reference and when to reinforce its own discoveries, while continuously upgrading what it imitates. We instantiate this paradigm as Active Group Relative Policy Optimization (Active-GRPO), realized through two coupled mechanisms: active imitate-reinforce and active referencing. The former performs imitation learning when the reference still outperforms the policy's own candidates, and shifts to self-improvement via reinforcement learning once the policy has generated molecules that surpass the reference. The latter continuously upgrades the reference itself by replacing it with the best policy-generated candidate discovered so far, progressively raising the imitation target and ensuring that reference guidance remains informative-rather than restrictive-throughout training. Across TOMG-Bench MOLOPT, Active-GRPO improves average SRxSim from 0.0959 for GRPO and 0.1665 for RePO to 0.1773 under matched three-seed evaluation, with statistically significant gains on LogP, MR, and QED.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!