2608.01666v1 Aug 03, 2026 cs.CL

스타일은 승리하고, 본질은 패배한다: 아이디어 생성 시 LLM을 평가자로 활용하는 문제점 진단

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

Fengxian Ji
Fengxian Ji
Citations: 20
h-index: 2
Zhuohan Xie
Zhuohan Xie
Citations: 327
h-index: 9
Min Peng
Min Peng
Citations: 476
h-index: 9
Fan Zhang
Fan Zhang
Citations: 16
h-index: 3
Qianqian Xie
Qianqian Xie
Citations: 443
h-index: 8
Jingpu Yang
Jingpu Yang
Citations: 16
h-index: 2
Zhexuan Cui
Zhexuan Cui
Citations: 7
h-index: 2
Xiuying Chen
Xiuying Chen
Citations: 143
h-index: 5
Yuke Li
Yuke Li
Citations: 73
h-index: 5
Juanfan Wu
Juanfan Wu
Citations: 0
h-index: 0
Yuansheng Xie
Yuansheng Xie
Citations: 3
h-index: 1

그러나 이러한 평가자들이 실제로 아이디어의 과학적 내용을 평가하는지, 아니면 표면적인 스타일 표현에 영향을 받는지 여부는 여전히 미결론적인 문제입니다. 이 질문에 답하기 위해, 우리는 LLM 기반 아이디어 평가에서 스타일 편향을 진단하고 완화하기 위한 통합된 세 가지 구성 요소 벤치마크인 SciStyleBench를 제안합니다. (i) 첫째, SciStyleStage는 고정된 과학적 내용에 대해 세 가지 환경(문맥 없음, 특정 분야 문맥, 개방형 도메인 검색 문맥)에서 통제된 스타일 변경을 적용하는 세 단계 평가 환경으로, 600개의 과학 아이디어와 15가지 스타일 변형을 포함하며, 각 환경당 9,000개의 평가 인스턴스를 제공합니다. (ii) 둘째, SciStyleMetrics는 Style Bias Index (SBI), Substance Recognition Rate (SRR), 그리고 Adversarial Win Rate (AWR)를 포함하는 일련의 정량적 지표로, 스타일 변형이 점수 안정성, 내용 식별 능력 및 순위 견고성에 미치는 영향을 분석합니다. (iii) 셋째, SciStyleExtractor는 스타일을 분리하여 과학적 내용을 평가하기 위한 즉시 사용 가능한 평가 모듈로, 스타일 조건부 평가 전에 스타일 유형과 편차를 예측하여, 스타일 인지력이 스타일 편향을 줄이는 데 도움이 되는지 평가할 수 있습니다. SciStyleBench에 대한 실험 결과, LLM 평가는 여전히 작성 스타일에 민감하며 과학적 내용을 구별하는 데 어려움을 겪는 것으로 나타났습니다. 반면, SciStyleExtractor는 SBI를 0.566에서 0.501로 줄이는 동시에 SRR과 AWR를 각각 0.504 및 0.554에서 0.759 및 0.899로 증가시켰습니다. 이러한 결과는 견고한 아이디어 평가가 과학적 내용에 대한 민감도를 희생하지 않고도 스타일 변형에 불변해야 함을 시사합니다. 전반적으로, SciStyleBench는 과학적 아이디어 평가에서 스타일 편향을 식별, 정량화 및 완화하기 위한 체계적인 프레임워크를 제공합니다.

Original Abstract

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias. Experiments on SciStyleBench show that direct LLM judges remain sensitive to writing style and struggle to distinguish scientific substance. In contrast, SciStyleExtractor reduces SBI from 0.566 to 0.501 while increasing SRR and AWR from 0.504 and 0.554 to 0.759 and 0.899, respectively. These results suggest that robust idea evaluation requires invariance to stylistic variation without sacrificing sensitivity to scientific substance. Overall, SciStyleBench provides a systematic framework for identifying, quantifying, and mitigating stylistic bias in scientific idea evaluation.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!