DIAGPaper: 멀티 에이전트 추론을 통한 과학 논문의 유효하고 구체적인 약점 진단
DIAGPaper: Diagnosing Valid and Specific Weaknesses in Scientific Papers via Multi-Agent Reasoning
단일 에이전트 또는 멀티 에이전트 LLM을 활용한 논문 약점 식별은 많은 관심을 받고 있지만, 기존 접근 방식은 중요한 한계를 가지고 있습니다. 많은 멀티 에이전트 시스템이 인간의 역할을 피상적으로 모방하여, 전문가들이 논문의 상호 보완적인 지적 측면을 평가하는 데 필요한 근본적인 기준을 놓치고 있습니다. 또한, 기존 방법은 식별된 약점이 유효하다고 암묵적으로 가정하며, 심사위원의 편향, 오해, 그리고 검토 품질을 검증하는 데 있어 저자의 반박이 갖는 중요한 역할을 간과합니다. 마지막으로, 대부분의 시스템은 사용자가 가장 중요한 문제를 우선적으로 파악할 수 있도록 순위를 매기지 않고 약점 목록을 그대로 제시합니다. 본 연구에서는 이러한 문제점을 해결하기 위해 세 가지 핵심 모듈이 긴밀하게 통합된 새로운 멀티 에이전트 프레임워크인 DIAGPaper를 제안합니다. 커스터마이저 모듈은 사용자가 정의한 검토 기준을 시뮬레이션하고, 기준별 전문성을 가진 여러 심사위원 에이전트를 생성합니다. 반박 모듈은 저자 에이전트를 도입하여 심사위원 에이전트와 체계적인 토론을 진행함으로써 제안된 약점을 검증하고 개선합니다. 우선순위 지정 모듈은 대규모의 인간 검토 사례로부터 학습하여 검증된 약점의 심각성을 평가하고, 사용자에게 가장 심각한 상위 K개의 약점을 제시합니다. AAAR 및 ReviewCritique라는 두 가지 벤치마크에서 실시한 실험 결과, DIAGPaper는 기존 방법보다 훨씬 우수한 성능을 보이며, 더 유효하고 논문에 특화된 약점을 제시하고, 사용자 중심적인 방식으로 우선순위를 매긴다는 것을 보여주었습니다.
Paper weakness identification using single-agent or multi-agent LLMs has attracted increasing attention, yet existing approaches exhibit key limitations. Many multi-agent systems simulate human roles at a surface level, missing the underlying criteria that lead experts to assess complementary intellectual aspects of a paper. Moreover, prior methods implicitly assume identified weaknesses are valid, ignoring reviewer bias, misunderstanding, and the critical role of author rebuttals in validating review quality. Finally, most systems output unranked weakness lists, rather than prioritizing the most consequential issues for users. In this work, we propose DIAGPaper, a novel multi-agent framework that addresses these challenges through three tightly integrated modules. The customizer module simulates human-defined review criteria and instantiates multiple reviewer agents with criterion-specific expertise. The rebuttal module introduces author agents that engage in structured debate with reviewer agents to validate and refine proposed weaknesses. The prioritizer module learns from large-scale human review practices to assess the severity of validated weaknesses and surfaces the top-K severest ones to users. Experiments on two benchmarks, AAAR and ReviewCritique, demonstrate that DIAGPaper substantially outperforms existing methods by producing more valid and more paper-specific weaknesses, while presenting them in a user-oriented, prioritized manner.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.