2607.27845v1 Jul 30, 2026 cs.CL

AutoSupervision: 근거 기반 수정 검증을 통한 과학 연구 워크플로우의 피드백 루프 완결

AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

Fenghua Ling
Fenghua Ling
Citations: 943
h-index: 17
Feng Liu
Feng Liu
Citations: 42
h-index: 4
Zijie Guo
Zijie Guo
Citations: 107
h-index: 5
Jiong Wang
Jiong Wang
Citations: 32
h-index: 4
Ben Fei
Ben Fei
Citations: 205
h-index: 9
Zixin Chen
Zixin Chen
Citations: 44
h-index: 3
Lei Bai
Lei Bai
Citations: 146
h-index: 6
Haobo Li
Haobo Li
Citations: 64
h-index: 5
Eunseo Jung
Eunseo Jung
Citations: 17
h-index: 2
Wenxiao Zhao
Wenxiao Zhao
Citations: 0
h-index: 0
Kaiyi Xu
Kaiyi Xu
Citations: 4
h-index: 1

최근 대규모 언어 모델(LLM)의 발전은 인공지능 시스템이 과학 연구 및 동료 심사를 지원할 수 있도록 했습니다. 그러나 신뢰성 있는 AI 기반 과학 워크플로우를 위한 필수적인 기능인, 심사위원 피드백이 의미 있고 증거에 기반한 논문 개선으로 이어지는지 확인하는 것은 아직 충분히 연구되지 않았습니다. 본 논문에서는 AutoSupervision을 소개합니다. 이는 과학 논문 수정 사항이 실제로 심사위원의 우려사항을 해결하는지, 근거를 통해 평가합니다. AutoSupervision은 투명한 동료 심사 기록을 자연스러운 지도 데이터 소스로 활용하며, 여기에는 심사위원 의견(과학적 우려사항), 저자의 답변(주장된 해결책), 수정된 논문(변경 사항에 대한 증거)이 포함됩니다. 심사위원 의견, 저자 답변 및 수정된 논문을 기반으로 모델은 심사위원의 우려사항을 파악하고, 이러한 우려사항이 해결되었는지 판단하며, 이를 뒷받침하는 논문의 증거를 식별해야 합니다. 우리는 56,000개의 Nature Communications 논문과 해당 동료 심사 기록을 사용하여 AutoSupervision을 구축했습니다. 그런 다음 LLM에 대한 실험, Ablation study 및 사례 연구를 수행했습니다. 우리의 결과는 LLM이 심사위원의 우려사항 파악에는 잘 작동하지만(GPT-5.5의 경우 0.754), 증거 기반 검증은 여전히 주요 병목 현상이며, 가장 성능이 좋은 모델도 0.501에 그친다는 것을 보여줍니다.

Original Abstract

Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence-supported manuscript improvements. We introduce AutoSupervision, which evaluates whether scientific manuscript revisions genuinely address reviewer concerns through grounded evidence. AutoSupervision leverages transparent peer-review records as a natural source of supervision, where reviewer comments specify scientific concerns, author responses describe claimed resolutions, and revised manuscripts provide evidence of changes. Given reviewer comments, author responses, and revised manuscripts, models must characterize reviewer concerns, determine whether concerns have been addressed, and identify supporting manuscript evidence. We construct AutoSupervision from 56,000 Nature Communications articles and corresponding review records. Then we conducted experiments on LLMs, the ablation study, and the case study. Our results show that while LLMs perform well in characterizing reviewer concerns, with GPT-5.5 achieving a score of 0.754, evidence-based verification remains the primary bottleneck, with the best-performing model reaching only 0.501.

1 Citations
0 Influential
8.5 Altmetric
43.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!