2603.15164v1 Mar 16, 2026 cs.CL

HindSight: 미래 영향력을 통한 연구 아이디어 생성 평가

HindSight: Evaluating Research Idea Generation via Future Impact

Bowen Jiang
Bowen Jiang
Citations: 7
h-index: 2

인공지능이 생성한 연구 아이디어를 평가하는 데 일반적으로 LLM(Large Language Model) 심사위원 또는 인간 패널이 사용되는데, 이는 주관적이며 실제 연구 영향력과는 동떨어진 경우가 많습니다. 본 연구에서는 미래의 연구 성과를 기준으로 아이디어의 품질을 평가하는 시간 분할 평가 프레임워크인 HindSight( extit{hs})를 소개합니다. extit{hs}는 생성된 아이디어를 실제 미래의 논문과 비교하여 인용 횟수 및 학술지 게재 여부 등을 기준으로 점수를 매깁니다. 특정 시점 $T$를 기준으로, 아이디어 생성 시스템은 $T$ 이전의 문헌만을 참조하도록 제한하고, 생성된 결과물을 $T$ 이후 30개월 동안 발표된 논문과 비교하여 평가합니다. 10개의 AI/ML 연구 분야에 대한 실험 결과, 놀라운 차이가 드러났습니다. LLM 심사위원은 검색 증강 방식과 일반적인 아이디어 생성 방식 사이에 유의미한 차이가 없다고 판단했습니다($p=0.584$). 반면, extit{hs}는 검색 증강 방식이 2.5배 더 높은 점수를 받는 아이디어를 생성한다는 것을 보여주었습니다($p<0.001$). 더욱이, extit{hs}의 점수는 LLM이 판단한 참신성과 extit{부정적} 상관관계를 보였습니다($ρ=-0.29$, $p<0.01$). 이는 LLM이 실제 연구로 이어지지 않는 참신하게 들리는 아이디어를 과대평가하는 경향이 있음을 시사합니다.

Original Abstract

Evaluating AI-generated research ideas typically relies on LLM judges or human panels -- both subjective and disconnected from actual research impact. We introduce \hs{}, a time-split evaluation framework that measures idea quality by matching generated ideas against real future publications and scoring them by citation impact and venue acceptance. Using a temporal cutoff~$T$, we restrict an idea generation system to pre-$T$ literature, then evaluate its outputs against papers published in the subsequent 30 months. Experiments across 10 AI/ML research topics reveal a striking disconnect: LLM-as-Judge finds no significant difference between retrieval-augmented and vanilla idea generation ($p{=}0.584$), while \hs{} shows the retrieval-augmented system produces 2.5$\times$ higher-scoring ideas ($p{<}0.001$). Moreover, \hs{} scores are \emph{negatively} correlated with LLM-judged novelty ($ρ{=}{-}0.29$, $p{<}0.01$), suggesting that LLMs systematically overvalue novel-sounding ideas that never materialize in real research.

1 Citations
0 Influential
1 Altmetric
6.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!