2608.08975v1 Aug 10, 2026 cs.CL

수사학적 요소가 인공지능 심사 시스템에 미치는 영향: 인공지능 기반 동료 평가에서 수사학적 민감성을 분석

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

Peng Shi
Peng Shi
Citations: 51
h-index: 4
Xirui Li
Xirui Li
Citations: 341
h-index: 6
Ming Li
Ming Li
Citations: 3
h-index: 1
Tianyi Zhou
Tianyi Zhou
Citations: 946
h-index: 11
Dawei Zhou
Dawei Zhou
Citations: 108
h-index: 6
Xinyue Zeng
Xinyue Zeng
Citations: 69
h-index: 4
Dianqi Li
Dianqi Li
Citations: 81
h-index: 4
Chenguang Wang
Chenguang Wang
Citations: 30
h-index: 2

대규모 언어 모델이 과학적 평가에 점점 더 많이 활용됨에 따라, 본 연구에서는 잠재적인 '보상 해킹' 형태인 '수사학적 선택이 보고된 과학적 내용이 동일할 때 인공지능 심사 결과에 미치는 영향', 그리고 이러한 효과가 평가 조건에 따라 어떻게 달라지는지를 조사합니다. 120개의 익명화된 ICLR 2026 제출 논문을 기반으로, 4,200개의 전체 논문 원고를 생성하고, 두 개의 대규모 언어 모델(LLM) 리라이터가 여섯 가지 수사학적 차원을 반대 방향으로 변환합니다. 다섯 개의 LLM 심사관이 표준 및 엄격한 프로토콜에 따라 결과물을 평가합니다. 또한, 공동 리라이팅, 재귀 리라이팅, 그리고 심사관 지향형 리라이팅을 테스트했습니다. 연구 결과는 수사학적 민감성이 균일하지 않고 구조화되어 있음을 보여줍니다. 증거 제시 방식과 새로운 주장 강조는 전체 평가에서 가장 큰 긍정-부정 대비를 나타내며, 범위 제시 방식은 그 다음으로 약한 수준의 영향을 미칩니다. 나머지 요소들은 상대적으로 작은 또는 불안정한 효과를 보입니다. 이러한 계층 구조는 인간이 평가한 품질 수준에 걸쳐 유지되지만, 점수 변화는 인공지능 심사관의 원래 점수에 크게 의존합니다. 낮은 점수는 상승하는 경향이 있고, 높은 점수는 하락하는 경향이 있으며, 방향성 대비는 중간 범위에서 가장 뚜렷하게 나타납니다. 더 복잡한 워크플로우가 항상 더 큰 개선을 가져오는 것은 아닙니다. 공동 리라이팅은 심각하게 리라이터에 의존적이며, 심사관 지향형 방식이 무작위 재평가보다 일관되게 우수한 결과를 내지 못합니다. 반복적인 리라이팅은 구성에 따라 감소하는 효과를 보입니다. 다양한 조건에서, 리라이터는 주로 반대 변형 간의 차이를 결정하고, 심사관은 점수 효과의 크기와 방향을 결정합니다. 엄격한 평가는 평균 OA(Overall Assessment) 점수를 1.36점 낮추지만, 수사학적 민감성에 일관된 변화를 가져오지는 않습니다. 이러한 연구 결과는 수사학적 표현이 인공지능 과학 심사에 미치는 영향을 명확히 하고, 내용 보존을 유지하면서도 과학적 글쓰기의 다양한 형태에 대해 강력한 평가 시스템을 개발하는 데 필요한 정보를 제공합니다.

Original Abstract

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!