2606.19749v1 Jun 18, 2026 cs.AI

에이전트 기반 검토 시스템의 성능 평가

Benchmarking Agentic Review Systems

Dang Nguyen
Dang Nguyen
University of Chicago
Citations: 26
h-index: 3
Chenhao Tan
Chenhao Tan
Citations: 232
h-index: 4
Wanqing Hao
Wanqing Hao
Citations: 0
h-index: 0
Yanai Elazar
Yanai Elazar
Bar-Ilan University
Citations: 5,743
h-index: 27

인공지능 지원 연구로 인해 피어 리뷰 시스템에 가해지는 압박을 완화하기 위한 새로운 유형의 에이전트 기반 검토 시스템이 등장하고 있지만, 이러한 시스템은 어떻게 평가해야 할지에 대한 명확성이 부족합니다. 본 연구에서는 OpenAIReview와 coarse라는 두 개의 오픈 소스 시스템, Reviewer3라는 독점 시스템, 그리고 제로샷 기준 모델을 포함하여 최첨단 및 효율적인 6개의 LLM을 사용하여 이들 시스템의 성능을 평가했습니다. 첫째, 인공지능 검토가 ICLR/NeurIPS 논문의 품질과 얼마나 일치하는지, 이를 외부 지표(인용 횟수, 합격 여부 등)를 통해 추정하여 분석했습니다. 모든 시스템은 쌍대 비교에서 우수한 성능을 보였으며, 가장 높은 정확도를 기록한 시스템은 OpenAIReview + GPT-5.5로 83.0%의 정확도를 달성했습니다. 둘째, 시스템이 알려진 정답을 기반으로 오류를 얼마나 잘 감지하는지 테스트하기 위해, 8개의 arXiv 주제 분류에 속하는 논문에 4가지 유형의 오류를 주입한 교란(perturbation) 벤치마크를 구축하고, 오류 탐지율을 측정했습니다. 가장 강력한 구성(OpenAIReview + GPT-5.5)은 주입된 오류의 71.6%를 탐지했으며, 이는 개선될 여지가 많음을 보여줍니다. 6개의 모델에서 탐지된 결과를 종합하면 83.3%의 재현율을 달성했습니다. 이는 서로 다른 모델이 서로 다른 오류를 탐지하며, 더 나은 설계가 성능 향상에 기여할 수 있음을 시사합니다. 또한, 실제 사용자를 대상으로 OpenAIReview를 공개적으로 배포하여 운영 결과를 분석한 결과, 댓글에 대한 긍정적인 반응이 압도적(1.44:1)이었으며, 가장 흔한 불만 사항은 오탐 및 사소한 지적사항에 대한 것이었습니다. 종합적으로, 최첨단 모델을 기반으로 하는 완전한 검토 시스템을 실제 연구 논문에 적용하여 평가한 결과, 인공지능 검토는 아직 개선의 여지가 있지만, 인간의 품질 판단과 잘 부합하며 중요한 오류를 탐지하고, 실제 사용자로부터 긍정적인 피드백을 받을 수 있음을 확인했습니다.

Original Abstract

A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but it is unclear how they should be evaluated. We evaluate two open-source systems (OpenAIReview and coarse), one proprietary system (Reviewer3), and a zero-shot baseline, across six LLMs spanning frontier and efficient models. First, we study whether AI reviews on ICLR/NeurIPS papers track with papers' quality as approximated by external signals such as citations and acceptance decisions. Every system performs above chance in pairwise accuracy, and the best is OpenAIReview + GPT-5.5 at 83.0%. Second, to test whether systems can catch errors with known ground truth, we construct a perturbation benchmark that injects four categories of errors into papers across eight arXiv subject classes and measure detection recall. The strongest configuration (OpenAIReview + GPT-5.5) catches 71.6% of injected errors, leaving substantial room for improvement. The union of detections across six models reaches 83.3% recall, suggesting different models detect different errors and better harness design can potentially increase performance. Beyond these benchmarks, we study a public deployment of OpenAIReview with real users. Votes on its comments skew positive at 1.44 to 1, and the most common complaints are about false positives and minor nitpicks. Together, by evaluating full review systems backed by state-of-the-art models on real research papers, we show that while AI reviews still have room for improvement, they can already track human quality judgments well, catch important errors, and earn positive feedback from real users.

0 Citations
0 Influential
13.5 Altmetric
67.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!