LLM을 활용한 연구의 참신성 평가: 한계에 대하여
On the Limits of LLM-as-Judge for Scientific Novelty Assessment
최근 LLM(Large Language Model, 거대 언어 모델)은 과학적 아이디어를 생성하고 평가하는 데 점점 더 많이 사용되고 있습니다. 이로 인해 연구의 참신성 평가는 중요한 문제가 되었습니다. 완전한 아이디어 평가는 종종 방법론 자체, 실현 가능성, 그리고 실제적인 유망성을 판단해야 하기 때문에 어렵습니다. 따라서 본 연구에서는 보다 근본적인 요소인 '연구 질문(RQ)'을 중심으로 분석합니다. 연구 질문 생성은 과학적 아이디어 발상의 전제 조건이며, 실제로 발표된 논문에서 다루어진 질문들과 비교할 수 있습니다. 우리는 최근 arXiv 논문을 기반으로 구축된 벤치마크인 RQ-Bench를 소개합니다. 각 논문에 대해 인용된 배경 지식, 해결되지 않은 문제점, 그리고 연구의 기여도를 바탕으로 작성자가 제시한 연구 질문을 재구성했습니다. 이러한 연구 질문은 해당 배경 지식에 대한 유일하게 올바른 질문이 아니지만, 참신성 판단을 위한 기준점으로 활용됩니다. 우리는 독립적인 LLM 평가, 비교 평가, 그리고 전문가 평가를 통해 모델이 생성한 연구 질문의 성능을 평가합니다. LLM 평가는 모델이 생성한 연구 질문을 일관되게 매우 새롭다고 평가하며, 이는 '참신성의 환상'을 야기합니다. 비교 평가에서는 이러한 선호도가 더욱 강화됩니다. 반면, 해당 분야 전문가들은 오히려 저자가 제시한 기준 연구 질문을 더 선호했습니다. 또한, 생성된 많은 연구 질문들이 범위가 좁거나 특정 자료에 의존적이라는 점을 발견했으며, LLM 평가는 이러한 측면을 종종 간과하는 경향이 있습니다. 종합적으로 볼 때, LLM 평가자와 전문가 간의 상반되는 참신성 판단은 LLM을 사용하여 연구 질문의 과학적 참신성을 평가하는 데 심각한 신뢰성 문제를 야기할 수 있습니다.
LLMs are increasingly used to generate and judge scientific ideas. This makes novelty evaluation a central problem. Full idea evaluation is difficult because it often requires judging a method, its feasibility, and its empirical promise. We therefore study a cleaner upstream object: the research question (RQ). RQ generation is a prerequisite for scientific ideation, and RQs can be compared against questions pursued in real papers. We introduce RQ-Bench, a benchmark built from recent arXiv papers. For each paper, we reconstruct author-anchored RQs from its cited background, gaps, and contributions. These RQs are not the only valid questions for the same background. They are author-anchored reference points for testing novelty judgments. We evaluate model-generated RQs with standalone LLM judging, comparative LLM judging, and human expert evaluation. LLM judges consistently rate model-generated RQs as highly novel, producing a novelty mirage; in comparative evaluations, this preference becomes even stronger. Domain experts, however, reach the opposite conclusion and prefer the author-anchored reference questions. We further find that many generated RQs are narrow or source-bound, a dimension that LLM judges often miss unless explicitly tested. Overall, the contradictory novelty evaluations between LLM judges and human experts raise a serious concern about the reliability of using LLMs to assess the scientific novelty of research questions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.