2606.23361v1 Jun 22, 2026 cs.LG

화학적 특성을 고려한 입력 검증 하에서의 분자 그래프 백도어 재고

Rethinking Molecular Graph Backdoors under Chemistry-aware Admission

Kok-Seng Wong
Kok-Seng Wong
Citations: 140
h-index: 4
Khoa D. Doan
Khoa D. Doan
Citations: 321
h-index: 8
Sze Jue Yang
Sze Jue Yang
Citations: 39
h-index: 4
Chee Seng Chan
Chee Seng Chan
Citations: 78
h-index: 3
Thinh Nguyen
Thinh Nguyen
Citations: 5
h-index: 1

분자 그래프 신경망(GNN)에 대한 백도어 공격은 일반적으로 추상적인 그래프 편집으로 평가되지만, 실제 분자 학습 파이프라인은 임의의 그래프로 훈련되지 않습니다. 분자 데이터는 먼저 구문 분석, 정제, 표준화 및 그래프-문자열 일관성 검사를 통과해야 합니다. 본 연구에서는 간과되어 온 이 입력 검증 단계를 ChemGuard라는 운영 프로토콜로 공식화하여, 제출된 분자 데이터가 실제 학습 파이프라인에 적합한지 테스트하고 기존 방어 기법을 보완합니다. ChemGuard는 분자 문자열이 정제 가능하고 해당 문자열에서 재구성된 그래프가 제출된 분자 그래프와 일치하는 경우에만 데이터를 허용합니다. 이러한 운영 관점에서, 많은 기존의 그래프 기반 백도어가 화학적으로 유효하지 않거나 표현 방식이 일치하지 않아 효과가 크게 저하되는 것을 알 수 있습니다. 그러나 입력 검증만으로는 분자 백도어를 완전히 차단하기에는 불충분하다는 점을 보여줍니다. 본 연구에서는 ChemBack이라는 입력 검증 단계를 고려한 분자 백도어 공격을 제안합니다. ChemBack은 화학적으로 타당한 모티프-앵커 연결을 구성하고, 입력 검증을 통과한 후보들을 클린 대상 클래스 분자와의 지문 기반 Tanimoto 유사성으로 순위를 매깁니다. ChemBack은 트리거 선택 시 모델에 의존하지 않으며, 분자 구조, 타겟 레이블, 지문 및 공개 유효성 검사만을 사용하고, 공격 대상 모델, 대체 GNN, 학습된 임베딩, 기울기, 로짓 또는 훈련 코드에 접근하지 않습니다. 다양한 분자 벤치마크, 검증 도구, 아키텍처 및 방어 기법에서 ChemBack은 입력 검증을 통과한 데이터에 대해 높은 공격 성공률을 달성하면서도 클린 데이터의 정확도를 유지합니다. 본 연구 결과는 양면적인 교훈을 제공합니다. 화학적 특성을 고려한 입력 검증은 많은 그래프 기반 백도어를 억제하지만, 화학적으로 유효하고 타겟 클래스에 맞는 분자 백도어는 여전히 실질적인 위협이라는 점입니다.

Original Abstract

Backdoor attacks on molecular graph neural networks (GNNs) are typically evaluated as abstract graph edits, but real molecular learning pipelines do not train on arbitrary graphs. Molecular records must first survive parsing, sanitization, canonicalization, and graph-string consistency checks. We formalize this overlooked admission stage as ChemGuard, an operational protocol for testing whether a submitted molecular record can enter a realistic learning pipeline, while complementing existing defenses. ChemGuard admits a record only when its molecular string is sanitizable and the graph reconstructed from that string matches the submitted molecular graph. Under this operational view, many existing graph-based backdoors lose much of their apparent efficacy because their poisons are chemically invalid or representation-inconsistent. We then show that admission checks alone are insufficient to rule out molecular backdoors. We propose ChemBack, an admission-aware molecular backdoor attack that constructs chemically feasible motif-anchor attachments and ranks admitted candidates by fingerprint-based Tanimoto similarity to clean target-class molecules. ChemBack is model-free during trigger selection, using molecular structures, target labels, fingerprints, and public validity checks, but no victim model, surrogate GNN, learned embedding, gradient, logit, or training-code access. Across molecular benchmarks, validators, architectures, and defenses, \textbf{ChemBack} achieves high attack success with fully admitted poisons while preserving clean accuracy. Our results reveal a two-sided lesson, chemistry-aware admission suppresses many graph-only backdoors, yet chemically valid and target-aligned molecular backdoors remain a practical threat.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!