ProtRLSearch: 강화 학습을 통해 훈련된 대규모 언어 모델 기반의 다중 라운드 다중 모드 단백질 검색 에이전트
ProtRLSearch: A Multi-Round Multimodal Protein Search Agent with Large Language Models Trained via Reinforcement Learning
의료 환경에서 발생하는 단백질 분석 작업은 종종 단백질 서열 제약 조건 하에서 정확한 추론을 요구하며, 질병 관련 변이의 기능 해석, 임상 연구를 위한 단백질 수준 분석, 그리고 유사한 시나리오를 포함합니다. 이러한 작업을 해결하기 위해, 단백질 관련 정보를 검색하는 검색 에이전트가 도입되어 질병 관련 변이 분석 및 단백질 기능 추론을 위한 의사 결정 지원을 제공합니다. 그러나 이러한 검색 에이전트는 주로 단일 라운드, 텍스트 기반 검색에만 제한되어 있어, 단백질 서열 정보를 다중 모드 입력으로 검색 의사 결정 과정에 통합하는 데 어려움이 있습니다. 또한, 최종 답변에만 초점을 맞춘 강화 학습(RL)의 감독 방식에 의존하기 때문에, 검색 과정에 대한 제약이 부족하여 키워드 선택 및 추론 방향의 오류를 적시에 식별하고 수정하기 어렵습니다. 이러한 제한 사항을 해결하기 위해, 우리는 다차원 보상을 기반으로 강화 학습을 통해 훈련된 다중 라운드 단백질 검색 에이전트인 ProtRLSearch를 제안합니다. ProtRLSearch는 실시간 검색 과정에서 단백질 서열과 텍스트를 다중 모드 입력으로 동시에 활용하여 고품질 보고서를 생성합니다. 모델이 단백질 서열 정보를 통합하고 실제 단백질 검색 환경에서 텍스트 기반 다중 모드 입력을 활용하는 능력을 평가하기 위해, 세 가지 난이도 수준으로 구성된 3,000개의 객관식 문제(MCQ)로 구성된 벤치마크인 ProtMCQs를 구축했습니다. 이 벤치마크는 단백질 기능 및 표현형 변화에 대한 서열 제약 추론부터 신호 전달 경로 및 조절 네트워크와 다차원 서열 특징을 통합하는 포괄적인 단백질 추론까지 다양한 단백질 관련 작업을 평가합니다.
Protein analysis tasks arising in healthcare settings often require accurate reasoning under protein sequence constraints, involving tasks such as functional interpretation of disease-related variants, protein-level analysis for clinical research, and similar scenarios. To address such tasks, search agents are introduced to search protein-related information, providing support for disease-related variant analysis and protein function reasoning in protein-centric inference. However, such search agents are mostly limited to single-round, text-only modality search, which prevents the protein sequence modality from being incorporated as a multimodal input into the search decision-making process. Meanwhile, their reliance on reinforcement learning (RL) supervision that focuses solely on the final answer results in a lack of search process constraints, making deviations in keyword selection and reasoning directions difficult to identify and correct in a timely manner. To address these limitations, we propose ProtRLSearch, a multi-round protein search agent trained with multi-dimensional reward based RL, which jointly leverages protein sequence and text as multimodal inputs during real-time search to produce high quality reports. To evaluate the ability of models to integrate protein sequence information and text-based multimodal inputs in realistic protein query settings, we construct ProtMCQs, a benchmark of 3,000 multiple choice questions (MCQs) organized into three difficulty levels. The benchmark evaluates protein query tasks that range from sequence constrained reasoning about protein function and phenotype changes to comprehensive protein reasoning that integrates multi-dimensional sequence features with signal pathways and regulatory networks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.