MolSight: 그래프 기반의 시각-언어 모델을 활용한 통합 화학 이미지 이해
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
분자 대규모 언어 모델(LLM)을 사용하여 분자 구조와 기능을 이해하는 것은 분자 설계 및 신약 개발과 같은 분야에서 새로운 트렌드로 부상하고 있습니다. 그러나 이러한 모델은 분자 구조의 시각적 표현을 완전히 포착하는 데 어려움을 겪으며, 잠재력을 제한합니다. 기존의 분자 시각-언어 모델(VLM)이 유망한 결과를 보여주지만, 여전히 구조 정렬 문제에 직면하고 있으며, 정확한 분자 이해를 위해서는 필요한 위상 모델링 기능이 부족합니다. 이러한 문제를 해결하기 위해, 본 논문에서는 VLM을 통해 분자 이미지의 이해를 향상시키기 위해 설계된 그래프 기반 시각-언어 모델 프레임워크인 MolSight를 제안합니다. MolSight는 화학 결합 연결 정보를 시각 토큰에 주입하는 분자 위상 모듈과, 시각적 특징을 화학 기호 의미와 정렬하는 분자 정지 모듈을 통합하여 구성됩니다. 실험 결과, MolSight는 다양한 화학 이미지 이해 작업에서 기존 VLM, 분자 LLM 및 특수 도구보다 훨씬 뛰어난 성능을 보이며, 분자 이미지 추론의 새로운 수준을 달성함을 보여줍니다.
Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in tasks such as molecular design and drug discovery. However, these models struggle to fully capture the visual representation of molecular structures, limiting their potential. While existing molecular vision-language models (VLMs) show promise, they still face challenges in structural alignment and lack the necessary topological modeling for accurate molecular understanding. To address this, we propose MolSight, a graph-aware vision-language model framework designed to enhance the understanding of molecular images by VLMs. MolSight integrates a Molecular Topology Module to inject chemical-bond adjacency information into vision tokens, and a Molecular Grounding Module to align visual features with chemical symbolic semantics. Our experiments demonstrate that MolSight significantly outperforms existing VLMs, molecular LLMs, and specialized tools across multiple chemical visual understanding tasks, achieving a new level of molecular image reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.