2607.00924v1 Jul 01, 2026 cs.AI

그래프 기반 강화 학습을 통한 추적 가능한 과학적 가설 생성: 개념 재조합을 활용

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

S. Pal
S. Pal
Citations: 110
h-index: 7
Shashwat Sourav
Shashwat Sourav
Citations: 48
h-index: 4
Tirthankar Ghosal
Tirthankar Ghosal
Citations: 411
h-index: 13
Markus J. Buehler
Markus J. Buehler
Citations: 109
h-index: 4

신소재 개발 속도를 높이기 위해서는 다단계, 분야에 특화된 추론을 통해 과학적으로 타당한 가설을 생성할 수 있는 AI 시스템이 필요합니다. 기존의 대규모 언어 모델은 종종 유창하지만 추적하기 어려운 답변을 생성하여 개방형 소재 설계 문제에 대한 최종 결과가 일관성 있는 중간 추론에 의해 뒷받침되는지 판단하기 어렵습니다. 본 연구에서는 Group Relative Policy Optimization (GRPO)으로 미세 조정된 그래프 기반 추론 모델인 Graph-PRefLexOR 패밀리를 개발했습니다. 이 모델은 메커니즘 탐색, 그래프 구성, 패턴 추출 및 가설 합성을 위한 명시적인 단계로 추론을 구성합니다. 이러한 설계는 신경망 언어 생성과 기호적 관계 구조를 연결하여 인과적 연결을 구축하고 검사하며 재사용할 수 있도록 합니다. 소재 과학 및 공학 문헌에서 가져온 100개의 개방형 질문에 대해 Graph-PRefLexOR은 해당 기본 모델보다 40~65%의 성능 향상을 보였으며, 특히 추론의 투명성 측면에서 가장 큰 개선을 달성했습니다. 임베딩 분석 결과, 제안된 모델은 더 넓은 의미 탐색 능력을 보여주었으며, 기준 모델에 비해 약 2~3배 더 높은 의미 다양성을 나타냈습니다. 또한, 의미적 역추적 및 레이어별 은닉 상태 분석을 통해 구조화된 추론과 최종 답변 간의 강한 일관성이 있음을 확인했습니다. 마지막으로, 테스트 시점의 그래프 확장 실험 결과, 추가적인 컴퓨팅 자원은 단순히 의미 범위를 넓히는 것이 아니라 제한된 의미 공간 내에서 장거리 개념 재조합을 증가시키는 데 주로 사용됩니다. 이러한 결과는 그래프 기반 강화 학습이 소재 설계 및 기타 과학 응용 분야에서 과학적 가설 생성을 위한 해석 가능한 AI 시스템 개발에 기여할 수 있음을 보여줍니다.

Original Abstract

Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models often produce fluent but weakly traceable responses to open-ended materials design problems, making it difficult to determine whether final answers are supported by coherent intermediate reasoning. We develop Graph-PRefLexOR, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization (GRPO) to organize reasoning into explicit phases for mechanism exploration, graph construction, pattern extraction, and hypothesis synthesis. This design links neural language generation with symbolic relational structure, enabling causal connections to be constructed, inspected, and reused. On 100 open-ended questions from materials science and mechanics literature, Graph-PRefLexOR achieves 40-65% improvements over corresponding base models, with the largest gains in reasoning traceability. Embedding analyses show broader semantic exploration and approximately 2-3 times greater semantic diversity than baselines. Semantic backtracking and layer-wise hidden-state analyses further show stronger alignment between structured reasoning and final answers. Finally, test-time graph expansion reveals that additional compute primarily increases long-range conceptual recombination within a bounded semantic space, rather than simply expanding semantic coverage. These results establish graph-native reinforcement learning as a pathway toward interpretable AI systems for scientific hypothesis generation in materials design and other scientific applications.

1 Citations
0 Influential
6.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!