RetroReasoner: 전략적 역합성 예측을 위한 추론 LLM
RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction
역합성 예측은 유기 합성의 핵심 과제로, 주어진 생성물 분자에 대한 반응 물질을 예측하는 것을 목표로 합니다. 전통적으로 화학자들은 가능한 결합 절단점을 선택하고 이에 해당하는 반응 물질을 유도하는데, 이는 시간 소모적이며 상당한 전문 지식을 요구합니다. 최근 분자 대규모 언어 모델(LLM)의 발전으로 일부 성과가 있었지만, 많은 방법들은 전략적 추론 없이 반응 물질을 예측하거나, 특정 반응 물질 선택으로 이어지는 논리적인 결합 절단 전략에 대한 명시적인 추론을 수행하는 대신 일반적인 생성물 분석만 수행합니다. 이러한 한계를 극복하기 위해, 우리는 화학자들의 전략적 사고를 활용하는 역합성 추론 모델인 RetroReasoner를 제안합니다. RetroReasoner는 지도 학습(SFT)과 강화 학습(RL)을 모두 사용하여 학습됩니다. SFT의 경우, 반응 물질 예측과 함께 구조화된 결합 절단 근거를 생성하는 프레임워크인 SyntheticRetro를 도입했습니다. RL의 경우, 예측된 반응 물질을 순방향 합성 모델을 통해 통과시켜 정확도를 보상으로 사용합니다. 예측된 생성물이 원래 입력 생성물과 일치하면 보상을 제공합니다. 실험 결과, RetroReasoner는 기존의 기본 모델보다 뛰어난 성능을 보일 뿐만 아니라, 더 광범위한 실행 가능한 반응 물질 제안을 생성하며, 특히 더 어려운 반응 사례를 처리하는 데 효과적임을 보여줍니다.
Retrosynthesis prediction is a core task in organic synthesis that aims to predict reactants for a given product molecule. Traditionally, chemists select a plausible bond disconnection and derive corresponding reactants, which is time-consuming and requires substantial expertise. While recent advancements in molecular large language models (LLMs) have made progress, many methods either predict reactants without strategic reasoning or conduct only a generic product analysis, rather than reason explicitly about bond-disconnection strategies that logically lead to the choice of specific reactants. To overcome these limitations, we propose RetroReasoner, a retrosynthetic reasoning model that leverages chemists' strategic thinking. RetroReasoner is trained using both supervised fine-tuning (SFT) and reinforcement learning (RL). For SFT, we introduce SyntheticRetro, a framework that generates structured disconnection rationales alongside reactant predictions. In the case of RL, we apply a round-trip accuracy as reward, where predicted reactants are passed through a forward synthesis model, and predictions are rewarded when the forward-predicted product matches the original input product. Experimental results show that RetroReasoner not only outperforms prior baselines but also generates a broader range of feasible reactant proposals, particularly in handling more challenging reaction instances.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.