2607.25672v1 Jul 28, 2026 astro-ph.IM

인공지능이 물리학, 천체물리학 및 우주론 분야의 과학 연구를 지원하는 능력: 문헌 검토 (I)

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

V. Krishnaraj
V. Krishnaraj
Citations: 0
h-index: 0
Kateryna Vovk
Kateryna Vovk
Citations: 5
h-index: 1
Kosuke Aizawa
Kosuke Aizawa
Citations: 0
h-index: 0
Adrian Bayer
Adrian Bayer
Citations: 21
h-index: 2
Linda Blot
Linda Blot
Citations: 368
h-index: 8
Jessica A Cowell
Jessica A Cowell
Citations: 78
h-index: 5
Suyog Garg
Suyog Garg
Citations: 3,385
h-index: 3
Jonathan Grée
Jonathan Grée
Citations: 4
h-index: 1
Anamaria Hell
Anamaria Hell
Citations: 187
h-index: 7
Ben Horowitz
Ben Horowitz
Citations: 0
h-index: 0
Masaya Ichikawa
Masaya Ichikawa
Citations: 0
h-index: 0
Kanyuni Iemoto
Kanyuni Iemoto
Citations: 0
h-index: 0
Keigo Kondo
Keigo Kondo
Citations: 0
h-index: 0
Zacharie Lorsin
Zacharie Lorsin
Citations: 0
h-index: 0
Kevin McCarthy
Kevin McCarthy
Citations: 4
h-index: 1
Jamie Robinson
Jamie Robinson
Citations: 0
h-index: 0
M. Ruiz-Granda
M. Ruiz-Granda
Citations: 125
h-index: 7
Ievgen Vovk
Ievgen Vovk
Citations: 181
h-index: 4
Ming-Hui Zhou
Ming-Hui Zhou
Citations: 0
h-index: 0
Jia Liu
Jia Liu
Citations: 0
h-index: 0
L. Thiele
L. Thiele
Citations: 478
h-index: 13

본 연구에서는 대규모 언어 모델(LLM)이 과학 연구를 위한 문헌 검토에 얼마나 효과적으로 활용될 수 있는지 조사합니다. 물리학, 천체물리학 및 우주론 분야의 전문가들이 설계한 8개의 연구 프로젝트를 대상으로 통제된 실험을 수행했습니다. 각 프로젝트는 명확하게 정의된 배경과 목표를 가지고 있으며, 인간 전문가와 AI 프롬프트 엔지니어는 동일한 문헌 검토 작업을 동시에 수행합니다. 인간이 선별한 관련 문헌과 2025년 중반 모델(ChatGPT-4o, ChatGPT Deep Research 및 Gemini)이 선별한 문헌을 비교했습니다. 그 결과, 인간과 AI가 선별한 참고문헌 간의 일치율은 매우 낮았습니다($<$6%), 이는 현재 AI 모델이 자체적으로 전문가 수준의 검색 능력을 갖추고 있지 않다는 것을 시사합니다. 하지만 AI는 인간의 문헌 검색 활동을 보완할 수 있는 잠재력이 있습니다. 또한, AI가 생성한 참고문헌 후보의 신뢰성과 완전성을 평가하고, 두 가지 유형의 환각 현상(fabrications: 존재하지 않는 논문에 대한 언급 및 metadata mismatches: 실제 논문에 대한 메타데이터 불일치)을 구별했습니다. 연구 결과, AI가 생성한 참고문헌 중 3%는 허구인 반면, 64%는 적어도 하나의 잘못된 필드(제목, 저자, 연도, 학술지, DOI 또는 링크)를 가진 실제 논문입니다. 이는 2025년 중반 모델에 대해 체계적인 검증이 필요하다는 것을 의미합니다. 그러나 2026년 모델인 ChatGPT Pro 5.5의 성능은 크게 향상되어, 단일 프로젝트 테스트에서 허구 또는 메타데이터 불일치가 전혀 발견되지 않았습니다.

Original Abstract

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!