분자 메시지 전달 신경망(MPNN)의 성능을 결정하는 요인은 무엇인가? 연산자 수준의 계승적 벤치마크
What drives performance in molecular MPNNs? An operator-level factorial benchmark
메시지 전달 신경망(MPNN)은 분자 특성 예측에 널리 사용되지만, 단일 구조로 구현되어 있어 특정 메시지 전달 연산자가 성능에 미치는 영향을 파악하기 어렵습니다. 본 연구에서는 2차원 분자 MPNN을 메시지 생성 초기화, 노드-엣지 병합 및 노드 업데이트 연산자의 세 가지 유형으로 분해하는 연산자 수준의 계승적 벤치마크를 제시합니다. 결과적으로 도출된 84개의 구성은 동일한 실험 설정과 통계 분석 프로토콜 하에서 MoleculeNet 데이터셋 10개에 대해 평가되었습니다. 이러한 엄격하게 제어된 환경에서, 성능 변동은 주로 업데이트 복잡성보다는 메시지 생성 방식과 관련이 깊습니다. 메시지 초기화는 회귀 및 분류 모두에서 가족 수준의 중요한 영향을 미치는 반면, 노드-엣지 병합은 연결 기반 혼합 방식을 사용하는 경우 회귀 분석에서 유의미한 가족 수준의 효과를 보입니다. 업데이트 패밀리는 어느 종단 지표에서도 통계적으로 유의미한 영향을 미치지 않습니다. 퀴네타존 분자에 대한 표현 연구 결과, 하다마드 게이팅보다 연결 기반 혼합 방식이 화학적으로 구별되는 헤테로 원자를 더 잘 구분하고 과도한 평활화를 방지하는 데 효과적임이 나타났습니다. 분류 및 회귀에 대해 별도로 선택된 대표적인 구성은 기존의 분자 그래프 신경망(GNN) 기준 모델과 경쟁력 있는 성능을 보이며, 벤치마크 데이터셋 10개 중 8개에서 수치적으로 가장 우수한 결과를 얻었습니다. 이러한 실증적 결과는 대표적인 노드-엣지 병합 및 업데이트 연산자에 대한 간결한 메커니즘 분석을 통해 해석되었습니다. 본 연구의 결과는 분자 MPNN 설계에 대한 경험적 지침을 제공하며, 모델 설계를 단일 구조 탐색에서 벗어나 메시지 전달 파이프라인 내 화학 정보가 어디에서 어떻게 입력되는지에 대한 목표 지향적인 평가로 전환하는 데 기여합니다.
Message-passing neural networks (MPNNs) are widely used for molecular property prediction, but their deployment as monolithic architectures makes it difficult to identify how specific message-passing operators affect performance. We present an operator-level factorial benchmark that decomposes 2D molecular MPNNs into the three families of message-seed initialization, node-edge fusion, and node update operators. The resulting 84 configurations are benchmarked on ten MoleculeNet datasets under a shared experimental setup and statistical analysis protocol. Across this controlled design, performance variation is associated primarily with message construction rather than update complexity. Message-seed initialization shows significant family-level effects for both regression and classification, node-edge fusion shows a significant family-level effect for regression with descriptive advantages for concatenation-based mixing, and the update family shows no statistically supported effect for either endpoint family. A representation probe into the Quinethazone molecule further demonstrates that concatenation-based mixing can better differentiate chemically distinct heteroatoms and withstand oversmoothing than Hadamard gating. Representative configurations selected separately for classification and regression recover competitive performance relative to established molecular graph neural network (GNN) baselines, ranking numerically best on eight of ten benchmark datasets. These empirical results are interpreted through concise mechanistic analyses of representative node-edge fusion and update operators. Our findings provide empirical design heuristics for molecular MPNNs by turning model design from a search over monolithic architectures into a targeted assessment of where and how chemical information enters the message-passing pipeline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.