향상된 부정 샘플링을 통한 지식 그래프 기반 모델 성능 향상
Boosting Knowledge Graph Foundation Models via Enhanced Negative Sampling
지식 그래프(KG)는 질의 응답 및 추천 시스템과 같은 다양한 하위 작업의 핵심적인 기반이 되었습니다. 그러나, 대부분의 지식 그래프는 불완전한 경우가 많습니다. 본 연구에서는 사전 학습에 사용된 것과는 다른 관계 어휘를 가진 새로운 지식 그래프에서 제로샷 지식 그래프 완성을 수행하기 위해, KG 기반 모델(KGFM)에 대한 관심이 높아지고 있습니다. 기존 KGFM은 종종 무작위로 생성된 부정 삼중항을 사용하여 훈련하는데, 이는 양의 삼중항의 머리 또는 꼬리 엔터티를 임의의 엔터티로 대체하여 구성됩니다. 그러나 이러한 부정 삼중항은 품질이 낮아 KGFM 훈련에 대한 약한 형태의 감독 신호를 제공합니다. 본 논문에서는 기존 KGFM을 향상시키기 위한 간단하면서도 효과적인 적응형 부정 샘플링 방법인 KMAS를 제안합니다. KMAS는 기존 KGFM의 관계 인코더에서 생성된 업데이트된 관계 임베딩을 사용하여 어려운 부정 삼중항을 구성합니다. 또한, KMAS는 훈련 과정 동안 KGFM의 진화하는 능력을 보다 효과적으로 반영하기 위해, 웜업 단계 이후 선형적으로 증가하고 그 후 선형적으로 감소하는 방식으로 전체 훈련 과정 동안 어려운 부정 삼중항의 비율을 동적으로 조정합니다. 총 44개의 데이터 세트를 사용하여 광범위한 실험을 수행했습니다. 실험 결과는 제안된 부정 샘플링 방법이 상당수의 최첨단 KGFM 성능을 향상시킬 수 있으며, 과도한 추가 시간이나 메모리 소모가 필요하지 않음을 보여줍니다.
Knowledge graphs (KGs) have become the core backbone of numerous downstream tasks such as question answering and recommender systems. However, despite all this, KGs are often very incomplete. To perform zero-shot knowledge graph completion in unseen KGs, which have different relational vocabularies from those used for pre-training, KG foundation models (KGFMs) receive a wide range of attention. Existing KGFMs often perform training using random negative triples, which are constructed by replacing the head or tail entity of a positive triple with a random entity. However, these negative triples are often constructed with limited quality, providing weak supervision for KGFM training. In this paper, we propose a simple yet effective adaptive negative sampling approach, KMAS, to enhance existing KGFMs. KMAS constructs hard negative triples through the updated relation embeddings generated from the existing KGFM's relation encoder. To further adaptively align with the evolving capability of the KGFM during the training process, KMAS adjusts the ratio of hard negative triples dynamically throughout the whole training process: after a warmup phrase, it increases the ratio linearly and then decreases linearly. Extensive experiments are conducted over 44 data sets. Experimental results demonstrate that our proposed negative sampling method can enhance many SOTA KGFMs without requiring excessive additional time or memory consumption.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.