OpenRTAG: 데이터 품질 저하 환경에서의 강력한 텍스트 속성 그래프 학습을 위한 종합적인 성능 평가 도구
OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation
텍스트 속성 그래프(Text-Attributed Graphs, TAG)는 관계 구조와 풍부한 노드 텍스트를 결합하는 중요한 그래프 데이터 형태입니다. 그러나 실제 환경의 TAG는 종종 불완전하며, 텍스트, 구조 및 레이블에서 발생하는 품질 문제로 인해 희소성, 노이즈 및 불균형을 나타냅니다. 이러한 요인들은 TAG 학습에 큰 영향을 미칠 수 있는 아홉 가지 대표적인 성능 저하 시나리오를 정의합니다. 기존 연구에서는 특정 완화 전략을 탐구했지만, 데이터 유형, 데이터셋, 작업 및 모델 유형에 따라 분산되어 있어 TAG의 강건성(robustness)에 대한 이해가 부족했습니다. 이러한 문제를 해결하기 위해, 텍스트 속성 그래프 학습을 위한 강건성 평가 도구인 OpenRTAG를 제안합니다. OpenRTAG는 TAG 품질 문제를 통일된 3x3 분류 체계로 구성하고, 아홉 개의 TAG 데이터셋과 세 가지 하위 작업에 대한 표준화된 평가를 지원합니다. 이를 통해 시나리오의 유효성과 모델 민감도를 체계적으로 평가하고, 기존 GNN(Graph Neural Networks), LLM-GNN(Large Language Model-based Graph Neural Networks) 및 대표적인 GFM(Graph Feature Matching) 모델을 비교하며, 특정 시나리오에 최적화된 기준 모델의 효과성, 효율성 및 강건성을 조사합니다. 또한 복합 성능 저하 시나리오에서의 모델 동작을 분석합니다. OpenRTAG는 현실적인 저품질 환경에서 TAG 학습의 강건성을 이해하기 위한 표준화된 테스트 환경을 제공합니다.
Text-attributed graphs (TAGs) are an important graph data form that combine relational structure with rich node text. However, real-world TAGs are often imperfect, with quality issues arising from text, structure, and labels, and typically manifesting as sparsity, noise, and imbalance. These dimensions define nine representative degradation scenarios that can substantially affect TAG learning. Although prior studies have explored specific mitigation strategies, existing evidence remains fragmented across degradation types, datasets, tasks, and model families, leaving TAG robustness insufficiently understood. To address this gap, we present OpenRTAG, a robustness benchmark for text-attributed graph learning. OpenRTAG organizes TAG quality issues into a unified 3 * 3 taxonomy and supports standardized evaluation across nine TAG datasets and three downstream tasks. It systematically evaluates scenario validity and model sensitivity, compares traditional GNNs, LLM-GNNs, and a representative GFM, investigates the effectiveness, efficiency, and robustness of scenario-matched baselines, and further examines model behavior under composite degradation scenarios. OpenRTAG provides a standardized testbed for understanding robustness in TAG learning under realistic low-quality settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.