AI 과학자 역량 평가를 위한 적대적이고 빠르게 변화하는 실제 환경 도메인: 성능 측정 기준
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
AI 과학자가 새로운 아이디어를 생성하는 능력을 평가하는 것은 매우 어렵습니다. 이 분야의 기존 벤치마크는 과학적 추론 및 연구 재현성을 평가하는 데 진전을 이루었지만, 종종 합성된 작업이나 사후 목표에 의존하며, 이는 사전 노출로 인해 왜곡될 수 있습니다. 우리는 전문가들이 독립적으로 관찰 가능한 결과를 생성하는 복잡하고 적대적인, 빠르게 변화하는 실제 환경 도메인이 AI 과학자의 역량, 즉 추론 능력, 창의성 및 가설 설정 능력을 평가하기 위한 실질적인 해결책을 제공할 수 있다고 가정합니다. 우리는 이 프레임워크를 구조적으로 다른 두 가지 도메인에 적용했습니다. 첫째는 2026 시즌의 자동차 디자인 컨셉에 대한 아이디어를 생성하는 Formula 1 (F1)이며, 실제 사전 시즈 혁신이 기준 데이터로 사용됩니다. 둘째는 Magic: The Gathering (MTG)이며, 모델은 최근 업데이트된 카드 풀에서 덱을 제안하고, 이를 19개의 Pro Tour (PT) 덱 목록과 비교하여 평가합니다. 두 도메인 모두에서 모델은 그럴듯한 결과를 생성하지만, 실제 전문가의 해결책과 일치하는 결과는 드뭅니다. F1에서는 가장 성능이 좋은 모델인 GPT-5.2가 실행 과정에서 제안된 166개의 아이디어 중 40개의 실제 혁신 중 10개를 찾아냈습니다. MTG에서는 Gemini 3 Flash 모델에서 생성된 최상의 덱이 세 번째로 높은 순위를 차지한 PT 덱에서 사용된 7개의 새로운 카드 중 5개를 복구했으며, 모든 108개의 덱 중에서 모델이 가장 자주 선택한 카드는 PT 덱에서 가장 널리 채택된 카드와 일치했습니다 (Spearman 상관 계수 ρ = 0.74, p = 0.0003). 이러한 결과는 AI 과학자에게 중요한 역량 격차가 아이디어 생성 자체가 아니라 필터링, 우선순위 결정 및 일관성 있는 창의성에 있다는 것을 시사합니다.
Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. We hypothesize that complex, adversarial, fast-moving real-world domains where expert practitioners independently generate observable outputs can provide a practical solution to fill this gap and evaluate the capabilities needed for AI scientists, including reasoning, novelty, and hypothesis formulation. We instantiate this framework in two structurally different domains, Formula 1 (F1), where models ideate around car design concepts for the 2026 season, and real pre-season innovations provide a ground truth, and Magic: The Gathering (MTG), where models propose decks from a recently updated card pool and are evaluated against 19 Pro Tour (PT) decklists. Across both domains, models produce plausible outputs, but few align with real-world expert solutions. In F1, the best model, GPT-5.2 matched 10 of 40 real innovations with 166 ideas proposed across runs. In MTG, the best deck from Gemini 3 Flash recovered 5 of 7 new-set cards from the third-place PT deck, and across all 108 decks, the cards models selected most often were also the cards most widely adopted by PT decks (Spearman $ρ= 0.74$, $p = 0.0003$). These results suggest that a key capability gap for AI scientists is not idea generation, but filtering, prioritization, and coherent novelty.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.