2603.21155v1 Mar 22, 2026 cs.AI

LLM이 그래프 학습을 속일 수 있을까? 텍스트 속성 그래프에 대한 범용 적대적 공격 탐색

Can LLMs Fool Graph Learning? Exploring Universal Adversarial Attacks on Text-Attributed Graphs

Xiang Ao
Xiang Ao
Citations: 1,383
h-index: 19
Kai Wu
Kai Wu
Citations: 42
h-index: 4
Yuling Wang
Yuling Wang
Citations: 181
h-index: 5
Pengfei Jiao
Pengfei Jiao
Citations: 31
h-index: 4
Xiao Wang
Xiao Wang
Citations: 9
h-index: 2
Dalin Zhang
Dalin Zhang
Citations: 22
h-index: 3
Zihui Chen
Zihui Chen
Citations: 8
h-index: 1

텍스트 속성 그래프(TAG)는 각 노드에 대한 풍부한 텍스트 의미론과 위상적 맥락을 통합하여 그래프 학습을 향상시킵니다. 이러한 표현력 향상은 동시에 텍스트 기반의 적대적 표면을 통해 그래프 학습에 새로운 취약점을 노출합니다. 최근 연구에서는 그래프 신경망(GNN) 및 사전 훈련된 언어 모델(PLM)과 같은 다양한 기반 모델을 활용하여 TAG에서 구조적 및 텍스트 정보를 모두 포착합니다. 이러한 다양성은 다음과 같은 중요한 질문을 제기합니다. TAG 모델의 보안을 평가하기 위해 다양한 아키텍처에 걸쳐 일반화되는 범용 적대적 공격을 어떻게 설계할 수 있을까요? 이 문제는 GNN과 PLM과 같이 서로 다른 기반 모델이 그래프 패턴을 인식하고 인코딩하는 방식의 현저한 차이점에서 비롯되며, 또한 많은 PLM이 API를 통해서만 접근할 수 있기 때문에 공격이 블랙박스 환경으로 제한되는 점도 문제입니다. 이러한 문제를 해결하기 위해, 본 연구에서는 대규모 언어 모델(LLM)의 일반적인 그래프 지식에 대한 이해를 깊이 파악하여 노드 위상과 텍스트 의미론을 동시에 변경하는 새로운 공격 프레임워크인 BadGraph를 제안합니다. 구체적으로, 본 연구에서는 그래프 사전 지식을 활용하여 모드 간에 정렬된 공격 단축 경로를 구축하는 타겟 영향력 검색 모듈을 설계하여 효율적인 LLM 기반의 교란 추론을 가능하게 합니다. 실험 결과, BadGraph는 GNN 및 LLM 기반 추론 시스템에 대해 보편적이고 효과적인 공격을 수행하며, 최대 76.3%의 성능 저하를 보였습니다. 또한 이론적 및 경험적 분석을 통해 BadGraph의 은밀하면서도 해석 가능한 특성을 확인했습니다.

Original Abstract

Text-attributed graphs (TAGs) enhance graph learning by integrating rich textual semantics and topological context for each node. While boosting expressiveness, they also expose new vulnerabilities in graph learning through text-based adversarial surfaces. Recent advances leverage diverse backbones, such as graph neural networks (GNNs) and pre-trained language models (PLMs), to capture both structural and textual information in TAGs. This diversity raises a key question: How can we design universal adversarial attacks that generalize across architectures to assess the security of TAG models? The challenge arises from the stark contrast in how different backbones-GNNs and PLMs-perceive and encode graph patterns, coupled with the fact that many PLMs are only accessible via APIs, limiting attacks to black-box settings. To address this, we propose BadGraph, a novel attack framework that deeply elicits large language models (LLMs) understanding of general graph knowledge to jointly perturb both node topology and textual semantics. Specifically, we design a target influencer retrieval module that leverages graph priors to construct cross-modally aligned attack shortcuts, thereby enabling efficient LLM-based perturbation reasoning. Experiments show that BadGraph achieves universal and effective attacks across GNN- and LLM-based reasoners, with up to a 76.3% performance drop, while theoretical and empirical analyses confirm its stealthy yet interpretable nature.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!