인간 연구 아이디어와 LLM(대규모 언어 모델) 연구 아이디어 간의 격차 측정
Measuring the Gap Between Human and LLM Research Ideas
LLM은 점점 더 많은 연구 아이디어 발상에 활용되고 있지만, 기존 평가 방법은 주로 개별 아이디어를 참신성, 실현 가능성 또는 전문가 선호도 측면에서 판단합니다. 본 연구에서는 다음과 같은 질문을 던집니다: 현재 LLM이 생성하는 아이디어는 인간 연구자들의 아이디어와 얼마나 동떨어져 있는가? 이 격차를 파악하기 위해, 고품질의 인간 연구 논문들을 기반으로 한 대규모 평가 프레임워크를 구축했습니다. 각 논문에 대해, 핵심 아이디어를 영감 받았을 가능성이 높은 관련 선행 연구들을 역추적합니다. 그런 다음, LLM에게 해당 논문의 제목과 요약 정보를 바탕으로 새로운 아이디어를 생성하도록 요청합니다. 우리는 각 아이디어를 기회 패턴과 연구 패러다임이라는 두 가지 축을 기준으로 분류하는 연구 경향 분류 체계를 도입하고, 이를 사용하여 인간과 LLM이 생성한 아이디어 간의 차이를 정량화합니다. 다양한 LLM에 의해 생성된 아이디어 세트에 대해 분석한 결과, 일관된 분포 차이가 나타났습니다. LLM이 생성한 아이디어는 주로 융합적인 기회와 합성 방법 주변에 집중되는 경향을 보이는 반면, 인간 연구 논문에 대한 참조 분포는 문제 정의 및 기여 방법 등 다양한 측면에 걸쳐 더 넓게 분포되어 있습니다. 이러한 결과는 강력한 LLM이 다양한 수준의 합리적인 아이디어를 생성할 수 있지만, 해당 범위가 여전히 인간의 연구 경향보다 좁고 체계적으로 다르다는 것을 시사합니다.
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research-taste taxonomy to profile each idea by its opportunity pattern and research paradigm, and use it to quantify the divergence between human and LLM ideas. Across idea sets generated by different LLMs, we observe a consistent distributional gap: LLM ideas are disproportionately concentrated around bridge-like opportunities and synthesis methods, whereas the human paper reference distribution spreads more broadly across ways of framing gaps and constructing contributions. This result suggests that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.