2605.27853v1 May 27, 2026 cs.AI

MolLingo: LLM 기반 과학 에이전트를 위한 분자 고유 표현

MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents

Heng Ji
Heng Ji
Citations: 911
h-index: 13
Thao Nguyen
Thao Nguyen
Citations: 21
h-index: 3

본 논문에서는 자동화된 분자 설계 프로세스를 모방하기 위해 설계된 멀티 에이전트 시스템인 MolLingo를 소개합니다. 기존의 LLM 기반 접근 방식은 외부 도구에 대한 접근 없이 독립적인 생성 모델로 작동하거나, 반복적이고 증거 기반의 추론을 수행하는 데 필요한 멀티 에이전트 조정 및 공유 메모리 기능을 갖추지 못했습니다. MolLingo는 공유 메모리 모듈을 통해 문헌 에이전트, 화학 에이전트 및 오케스트레이터를 조정하며, 각 에이전트는 도메인별 특화된 도구를 갖추고 있습니다. 효과적인 분자 추론을 위해, 우리는 합성 가능성을 고려한 분자 파편화 방법인 BRICS 기반 파편 열거(BFE)를 도입했습니다. 이 방법은 분자를 화학적으로 의미 있는 구성 요소로 분해하며, 각 구성 요소는 블록 기반 SMILES 표현과 일반적인 화학 명칭으로 표현됩니다. 이러한 표현 방식은 분자 구조와 LLM의 의미 공간을 연결하여, 원시 SMILES만으로는 어려운 블록 수준의 추론 및 편집을 가능하게 합니다. 초기 단계 치료제 설계 사례 연구를 통해, MolLingo는 분자 도킹으로부터 얻은 결합 부위 기하학적 정보 및 잔기 수준의 단백질 맥락을 활용하여 화학 에이전트의 추론을 뒷받침하고, 분자의 표적 결합력을 최적화합니다. 네 가지 벤치마크 테스트에서 MolLingo는 GPT-5.4를 포함한 선도적인 LLM 및 특수 기준 모델보다 일관되게 우수한 성능을 보였습니다. 특히 도킹 점수가 4배 향상되었으며, 다양한 LLM 백본에서 약물 특성 최적화 효과가 꾸준히 나타났고, TOMG-Bench에서는 선도적인 LLM 및 강화 학습 기반 최적화 방법인 RePO를 능가하는 최고 수준의 결과를 달성했습니다. 이러한 결과는 LLM이 화학적으로 의미 있는 표현과 생물학적으로 근거한 구조적 맥락을 통해 안내받을 때 이미 강력한 분자 설계 도구로 활용될 수 있음을 시사합니다. 코드: https://anonymous.4open.science/status/MolLingo-7450

Original Abstract

We present MolLingo, a multi-agent system that emulates the reasoning process of a chemist to automate molecular design. Existing LLM-based approaches either operate as standalone generative models without access to external tools or lack the multi-agent coordination and shared memory needed for iterative, evidence-driven reasoning across the molecular design pipeline. MolLingo addresses this by coordinating a Literature Agent, a Chemist Agent, and an Orchestrator through a shared memory module, with each agent equipped with domain-specific tools. To enable effective molecular reasoning, we introduce BRICS-based Fragment Enumeration (BFE), a synthesis-aware molecular fragmentation method that decomposes molecules into chemically meaningful building blocks represented as block-based SMILES paired with common chemical names. This representation bridges molecular structure and LLM semantic space, enabling block-level reasoning and editing that is difficult with raw SMILES alone. As a case study in early-stage therapeutic design, MolLingo further grounds the Chemist Agent's reasoning in binding site geometry and residue-level protein context derived from molecular docking to optimize molecules for stronger target binding. Across four benchmarks, MolLingo consistently outperforms frontier LLMs and specialized baselines, including a fourfold docking score improvement over GPT-5.4 despite using the same underlying model, consistent drug property optimization gains across multiple LLM backbones, and state-of-the-art results on TOMG-Bench, surpassing both frontier LLMs and the RL-based optimization method RePO. Our results suggest that LLMs are already capable molecular design assistants when guided through chemically meaningful representations and biologically grounded structural context. Code is available at: https://anonymous.4open.science/status/MolLingo-7450.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!