압축 지식 그래프 가설: 과학적 가설 생성에 어떤 그래프 사실이 중요한가?
The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?
지식 그래프(KG)는 언어 모델에 구조화된 과학적 맥락을 제공할 수 있지만, 실제로 생성된 가설에 영향을 미치는 그래프 사실은 무엇인지 불분명합니다. 본 연구에서는 Mistral-7B, Llama-3.1-70B, Gemini 2.5 Flash 모델을 사용하여 배터리 소재 관련 지식 그래프 기반 가설 생성을 연구했습니다. 우리는 밀도, 온톨로지 풍부성, 위상 구조 및 제어 구조를 변경하여 로컬 KG를 수정하고, 제공된 그래프와 고정 참조 메트릭을 모두 사용하여 결과를 평가했습니다. 모델 전반적으로, KG의 유용성은 선택적이며 모델에 따라 다릅니다: 그래프 맥락은 출력에 영향을 미치지만, KG가 없는 경우에도 모델의 사전 지식을 통해 상당한 양의 그래프 정보가 복구됩니다. 종종, 전체 KG와 유사한 결과를 보이는 상위 k개의 압축된 부분 그래프가 발견되며, 여기에는 주장된 결과(claimed-outcome) 트리플이 제외되는 경우도 포함됩니다. 동시에, 이러한 압축은 특정 의미적 순위 규칙에만 국한되지 않으며, 무작위 및 위상 기반의 부분 집합 또한 상당한 정보를 복구할 수 있습니다. 이러한 결과는 다음과 같은 '압축 지식 그래프' 가설을 뒷받침합니다: 유용한 KG 정보는 종종 전체 로컬 그래프를 필요로 하기보다는, 작고 과학적으로 구조화된 부분 그래프에서 복구될 수 있는 경향이 있습니다.
Knowledge graphs (KGs) can provide structured scientific context to language models, but it remains unclear which graph facts actually shape the generated hypotheses. We study KG-guided hypothesis generation for battery materials across Mistral-7B, Llama-3.1-70B, and Gemini 2.5 Flash. We perturb local KGs by varying density, ontology richness, topology, and control structure, and evaluate outputs with both provided-graph and fixed-reference metrics. Across models, KG utility is selective and model-dependent: graph context changes outputs, but no-KG outputs also recover substantial graph content from model priors. Compact top-k subgraphs often approximate full-KG behavior, including when claimed-outcome triples are held out. At the same time, compression is not unique to one semantic ranking rule, random and topology-based subsets can also recover much of the signal. These results support a redundancy-aware Compressive KG hypothesis: useful KG signal is often recoverable from compact, scientifically structured subgraphs rather than requiring the full local graph.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.