2602.06707v1 Feb 06, 2026 cs.AI

지식 그래프 생성을 위한 자기회귀 모델

Autoregressive Models for Knowledge Graph Generation

Thiviyan Thanapalasingam
Thiviyan Thanapalasingam
Citations: 529
h-index: 8
Antonis Vozikis
Antonis Vozikis
Citations: 0
h-index: 0
Peter Bloem
Peter Bloem
Citations: 17
h-index: 3
P. Groth
P. Groth
Citations: 49
h-index: 4

지식 그래프(KG) 생성은 모델이 도메인 유효성 제약 조건을 유지하면서 트리플 간의 복잡한 의미론적 의존성을 학습할 것을 요구한다. 트리플의 점수를 독립적으로 매기는 링크 예측과 달리, 생성 모델은 의미적으로 일관된 구조를 만들기 위해 전체 부분 그래프에 걸친 상호 의존성을 포착해야 한다. 우리는 그래프를 (head, relation, tail) 트리플의 시퀀스로 간주하여 KG를 생성하는 자기회귀 모델군인 ARK(Auto-Regressive Knowledge Graph Generation)를 제안한다. ARK는 명시적인 규칙 지도 없이 데이터로부터 타입 일관성, 시간적 유효성, 관계 패턴을 포함한 암시적 의미론적 제약 조건을 직접 학습한다. IntelliGraphs 벤치마크에서 우리 모델은 훈련 중에 보지 못한 새로운 그래프를 생성하면서 다양한 데이터셋에 걸쳐 89.2%에서 100.0%의 의미론적 타당성을 달성했다. 또한 우리는 학습된 잠재 표현을 통해 제어된 생성을 가능하게 하고, 비조건부 샘플링과 부분 그래프로부터의 조건부 완성을 모두 지원하는 ARK의 변분 확장 모델인 SAIL을 소개한다. 우리의 분석에 따르면 KG 생성에는 아키텍처의 깊이보다 모델 용량(은닉 차원 64 이상)이 더 중요하며, 순환(recurrent) 아키텍처가 상당한 계산 효율성을 제공하면서도 트랜스포머 기반 대안과 대등한 타당성을 달성함을 보여준다. 이러한 결과는 자기회귀 모델이 지식 베이스 완성 및 질의 응답에서의 실질적인 응용과 함께 KG 생성을 위한 효과적인 프레임워크를 제공함을 입증한다.

Original Abstract

Knowledge Graph (KG) generation requires models to learn complex semantic dependencies between triples while maintaining domain validity constraints. Unlike link prediction, which scores triples independently, generative models must capture interdependencies across entire subgraphs to produce semantically coherent structures. We present ARK (Auto-Regressive Knowledge Graph Generation), a family of autoregressive models that generate KGs by treating graphs as sequences of (head, relation, tail) triples. ARK learns implicit semantic constraints directly from data, including type consistency, temporal validity, and relational patterns, without explicit rule supervision. On the IntelliGraphs benchmark, our models achieve 89.2% to 100.0% semantic validity across diverse datasets while generating novel graphs not seen during training. We also introduce SAIL, a variational extension of ARK that enables controlled generation through learned latent representations, supporting both unconditional sampling and conditional completion from partial graphs. Our analysis reveals that model capacity (hidden dimensionality >= 64) is more critical than architectural depth for KG generation, with recurrent architectures achieving comparable validity to transformer-based alternatives while offering substantial computational efficiency. These results demonstrate that autoregressive models provide an effective framework for KG generation, with practical applications in knowledge base completion and query answering.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!