HindSight: 미래 영향력을 통한 연구 아이디어 생성 평가
HindSight: Evaluating Research Idea Generation via Future Impact
인공지능이 생성한 연구 아이디어를 평가하는 데 일반적으로 LLM(Large Language Model) 심사위원 또는 인간 패널이 사용되는데, 이는 주관적이며 실제 연구 영향력과는 동떨어진 경우가 많습니다. 본 연구에서는 미래의 연구 성과를 기준으로 아이디어의 품질을 평가하는 시간 분할 평가 프레임워크인 HindSight( extit{hs})를 소개합니다. extit{hs}는 생성된 아이디어를 실제 미래의 논문과 비교하여 인용 횟수 및 학술지 게재 여부 등을 기준으로 점수를 매깁니다. 특정 시점 $T$를 기준으로, 아이디어 생성 시스템은 $T$ 이전의 문헌만을 참조하도록 제한하고, 생성된 결과물을 $T$ 이후 30개월 동안 발표된 논문과 비교하여 평가합니다. 10개의 AI/ML 연구 분야에 대한 실험 결과, 놀라운 차이가 드러났습니다. LLM 심사위원은 검색 증강 방식과 일반적인 아이디어 생성 방식 사이에 유의미한 차이가 없다고 판단했습니다($p=0.584$). 반면, extit{hs}는 검색 증강 방식이 2.5배 더 높은 점수를 받는 아이디어를 생성한다는 것을 보여주었습니다($p<0.001$). 더욱이, extit{hs}의 점수는 LLM이 판단한 참신성과 extit{부정적} 상관관계를 보였습니다($ρ=-0.29$, $p<0.01$). 이는 LLM이 실제 연구로 이어지지 않는 참신하게 들리는 아이디어를 과대평가하는 경향이 있음을 시사합니다.
Evaluating AI-generated research ideas typically relies on LLM judges or human panels -- both subjective and disconnected from actual research impact. We introduce \hs{}, a time-split evaluation framework that measures idea quality by matching generated ideas against real future publications and scoring them by citation impact and venue acceptance. Using a temporal cutoff~$T$, we restrict an idea generation system to pre-$T$ literature, then evaluate its outputs against papers published in the subsequent 30 months. Experiments across 10 AI/ML research topics reveal a striking disconnect: LLM-as-Judge finds no significant difference between retrieval-augmented and vanilla idea generation ($p{=}0.584$), while \hs{} shows the retrieval-augmented system produces 2.5$\times$ higher-scoring ideas ($p{<}0.001$). Moreover, \hs{} scores are \emph{negatively} correlated with LLM-judged novelty ($ρ{=}{-}0.29$, $p{<}0.01$), suggesting that LLMs systematically overvalue novel-sounding ideas that never materialize in real research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.