LLM 기반 검색을 통한 분자 교란 반응 예측
LLM-Guided Retrieval for Prediction of Molecular Perturbation Responses
세포주에 대한 저분자 화합물의 유전체 반응 예측은 신약 개발의 핵심이지만, 모든 약물-세포 조합에 대한 광범위한 분석은 현실적으로 불가능합니다. 본 연구에서는 분자 교란 예측을 '검색 및 집계' 방식으로 정의하고, 측정되지 않은 특정 약물에 대한 세포주의 반응을 추정하기 위해, 관련성이 높은 소수의 화합물의 측정된 반응 결과를 종합하는 방법을 제시합니다. 우리는 LLM 기반 검색 (LLM-Guided Retrieval, LGR)이라는 새로운 접근 방식을 제안하는데, 이는 대규모 언어 모델(LLM)을 사용하여 대상 세포주에서 프로파일링된 후보 약물들 중에서 가장 적절한 'neighbor drug'를 순위화하는 방식입니다. 이후, 선택된 약물들의 관찰된 유전자 발현 변화량을 고정된 평균 집계기로 결합하여 예측값을 생성합니다. 우리는 Tahoe-100M 단일 세포 교란 데이터셋을 사용하여 새로운 약물, 새로운 세포주, 그리고 개방형 환경에서 LGR의 성능을 평가했습니다. 실험 결과, LGR은 기존의 단순 평균 방법, ChemCPA, 그리고 화학 기반 kNN 모델보다 일관적으로 우수한 성능을 보였으며, 특히 새로운 세포주에 대한 일반화 성능이 뛰어나 상관관계가 높고 오차가 적었습니다. 다양한 환경에서 LGR은 유전자 조절 방향(sign) 정확도를 향상시켜, 크기 기반 지표가 유사하더라도 생물학적으로 의미 있는 교란 효과를 더 잘 반영하는 것을 보여줍니다. 이러한 결과는 분자 교란 예측의 성능이 복잡한 예측 모델보다는 검색 품질에 크게 의존하며, LLM이 제약된 검색 모듈로 사용될 때 유용한 생물학적 사전 지식을 제공할 수 있음을 시사합니다.
Predicting transcriptomic responses to small-molecule perturbations across cell lines is central to drug discovery, but exhaustive profiling of drug-cell combinations is infeasible. We frame molecular perturbation prediction as retrieve-and-aggregate: approximate an unmeasured drug's response in a cell line by aggregating measured responses of a small set of biologically related compounds. We propose LLM-Guided Retrieval (LGR), where a large language model (LLM) ranks candidate neighbor drugs (restricted to those profiled in the target cell line); after which a fixed mean aggregator combines their observed expression deltas to form the prediction. We evaluate on the Tahoe-100M single-cell perturbation atlas under unseen-drug, unseen-cell-line, and open-world regimes. LGR consistently improves over drug mean, ChemCPA, and chemistry-based kNN baselines, with the strongest gains for unseen cell-line generalization, where it achieves higher correlation and lower error than mean baselines. Across settings, LGR improves directional (sign) accuracy of gene regulation, indicating better recovery of biologically meaningful perturbation effects even when magnitude-based metrics are similar. These results suggest that retrieval quality, rather than predictor complexity, is a key driver of zero-shot molecular perturbation prediction, and that LLMs can provide a useful biological prior when used as constrained retrieval modules.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.