아이디어는 유전체를 갖는다: 과학적 계보 추론 및 계보 기반 아이디어 생성 성능 평가
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
과학적 아이디어는 흔히 백지 상태에서 시작하지 않습니다. 오히려 생물학적 유전체와 마찬가지로, 기존 연구의 메커니즘을 상속받고, 알려진 한계를 개선하며, 이전 작업의 요소들을 재조합합니다. 현재의 성능 평가 기준은 인공지능 시스템이 이러한 상속 구조를 따를 수 있는지에 대해 아직 많은 것을 보여주지 못하고 있습니다. 본 논문에서는 과학적 계보 추론 및 계보 기반 아이디어 생성을 위한 벤치마크인 IdeaGene-Bench (IG-Bench)를 제시합니다. IG-Bench는 IdeaGene 프레임워크를 중심으로 구성되어 있으며, 각 논문 또는 제안은 최소한의, 유형화된, 증거 기반의 Idea Genome 객체 집합으로 표현됩니다. 또한 GenomeDiff는 이러한 객체들을 정렬하여 상속, 돌연변이, 소실, 외부 도입 및 새로운 삽입을 여섯 가지 운영적 진화 동역학 하에서 기록합니다. 벤치마크에는 10개의 과학 분야에 걸쳐 1,961개의 이상적인 계보 추적 정보, 1,085개의 큐레이션된 Idea Genome 객체 및 920개의 쌍별 GenomeDiff 레코드가 포함되어 있습니다. 이 벤치마크는 두 가지 평가를 지원합니다. IG-Exam (42가지 작업 유형, 1,029개 인스턴스)은 Idea Genome 추상화, 계보 추적, 진화적 추론 및 계보 검증을 통한 완전한 형태의 계보 추론 능력을 테스트합니다. IG-Arena는 Population-Evolution Score(PES)를 사용하여 아이디어 생성 성능을 평가하며, 제안이 주어진 계보 집단의 일관된 후손으로 삽입될 수 있는지 여부를 판단합니다. 즉, 적절한 Idea Genome 객체를 상속받고, 주변 연구와 의미 있게 달라져야 하며, 향후 연구에 가치를 제공해야 합니다. 14개의 LLM 기반 시스템을 대상으로 실시한 실험 결과, 구성 요소 결합의 어려움이 드러났습니다. 가장 성능이 좋은 시스템도 계보 추론에서 정확도가 27.3%에 불과했으며, 구조화된 계보 정보는 모든 참가자에게 균일하게 도움이 되는 것이 아니라 오히려 시스템 순위를 재조정하는 경향을 보였습니다.
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.