2607.25881v1 Jul 28, 2026 cs.CL

AI가 물리학, 천체물리학 및 우주론 연구의 과학적 연구 지원에 기여하는 능력 II: 프로젝트 계획 수립 및 제안서 평가

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

Jia Liu
Jia Liu
Citations: 39
h-index: 3
V. Krishnaraj
V. Krishnaraj
Citations: 0
h-index: 0
Kateryna Vovk
Kateryna Vovk
Citations: 5
h-index: 1
Kosuke Aizawa
Kosuke Aizawa
Citations: 0
h-index: 0
Adrian Bayer
Adrian Bayer
Citations: 21
h-index: 2
Linda Blot
Linda Blot
Citations: 368
h-index: 8
Jessica A Cowell
Jessica A Cowell
Citations: 78
h-index: 5
Suyog Garg
Suyog Garg
Citations: 3,385
h-index: 3
Jonathan Grée
Jonathan Grée
Citations: 4
h-index: 1
Anamaria Hell
Anamaria Hell
Citations: 187
h-index: 7
Ben Horowitz
Ben Horowitz
Citations: 0
h-index: 0
Masaya Ichikawa
Masaya Ichikawa
Citations: 0
h-index: 0
Kanyuni Iemoto
Kanyuni Iemoto
Citations: 0
h-index: 0
Keigo Kondo
Keigo Kondo
Citations: 0
h-index: 0
Zacharie Lorsin
Zacharie Lorsin
Citations: 0
h-index: 0
Kevin McCarthy
Kevin McCarthy
Citations: 4
h-index: 1
Jamie Robinson
Jamie Robinson
Citations: 0
h-index: 0
M. Ruiz-Granda
M. Ruiz-Granda
Citations: 125
h-index: 7
L. Thiele
L. Thiele
Citations: 33
h-index: 2
Ievgen Vovk
Ievgen Vovk
Citations: 181
h-index: 4
Ming-Hui Zhou
Ming-Hui Zhou
Citations: 0
h-index: 0

본 연구에서는 대규모 언어 모델(LLM)이 과학 프로젝트 계획 수립 및 제안서 평가를 얼마나 효과적으로 지원할 수 있는지 조사합니다. 인간 연구원과 세 개의 최신 LLM(ChatGPT, Claude, DeepSeek; 2025년 중반 모델, 기본 도구 액세스 사용)을 사용하여 물리학, 천체물리학 및 우주론 분야의 전문가가 설계한 8개의 연구 프로젝트에 대해 각각 1페이지 분량의 프로젝트 계획을 독립적으로 작성했습니다. 이렇게 생성된 32개의 제안서는 4명의 인간 평가자와 두 개의 최신 LLM(Claude Opus 4.8, ChatGPT Pro 5.5)이 4가지 측면으로 구성된 평가 기준에 따라 익명으로 평가했습니다. 또한 평가자는 각 제안서가 인간 또는 AI에 의해 작성되었는지 식별하도록 요청받았습니다. 인간 평가자는 전반적으로 인간이 작성한 제안서와 AI가 작성한 제안서를 유사하게 평가했지만, 두 LLM 평가자는 AI가 작성한 제안서를 인간이 작성한 제안서보다 평균 1점 정도 더 높게 평가했습니다(5점 만점). 인간 평가자는 각각 72% 및 79%의 정확도로 인간과 AI가 작성한 제안서를 식별했으며, 두 LLM 평가자는 모든 32개의 제안서를 100% 정확하게 분류했습니다. 이러한 결과는 현재 LLM이 인간이 작성한 것과 유사한 수준의 프로젝트 계획을 생성할 수 있지만, AI 평가자는 AI가 생성한 제안서에 대해 체계적으로 더 높은 선호도를 보인다는 것을 시사합니다. 본 연구 결과는 제안 준비 및 평가 과정에서 LLM을 광범위하게 사용하는 데 주의가 필요하다는 점을 강조합니다.

Original Abstract

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!