2605.26662v1 May 26, 2026 cs.CL

AI 평가가 인식에 편향을 초래할 수 있습니다: 학술 논문 해석 시 맥락의 중요성

AI evaluation may bias perceptions: The importance of context in interpreting academic writing

Shanglin Wu
Shanglin Wu
Citations: 27
h-index: 2
Randol Yao
Randol Yao
Citations: 0
h-index: 0

본 연구는 과학 논문에 사용된 AI 기술의 활용도를 추정하는 과정에서, 국가 및 분야 간의 맥락적 차이를 고려하지 않을 때 발생할 수 있는 편향 현상을 분석합니다. Dimensions 데이터베이스의 대규모 저널 출판 데이터를 사용하여, 인간이 작성한 초록과 LLM(대규모 언어 모델)으로 재작성된 초록 간의 차이를 기반으로 AI 유사도 벤치마크를 구축했습니다. 연구 결과, 통합된 벤치마크는 기존의 스타일적 다양성을 AI 생성 텍스트와 혼동시켜, LLM 등장 이전의 논문에서도 국가 및 분야별 그룹에 상당한 왜곡을 초래하는 것으로 나타났습니다. 반면, 국가 및 분야별로 특화된 벤치마크는 이러한 왜곡을 줄이고 비교를 위한 보다 신뢰할 수 있는 기준점을 제공합니다. 본 연구에서 제안된 방법론을 2025년의 논문에 적용한 결과, 통합된 벤치마크는 특정 국가 및 분야에서는 AI 사용량을 과도하게 추정하는 반면, 다른 국가 및 분야에서는 과소평가하는 경향이 있었습니다. 이러한 결과는 과학 분야에서 AI 활용도를 정확하고 공정하게 평가하기 위해서는 맥락을 고려한 측정 방법의 중요성을 강조합니다.

Original Abstract

This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct AI-likeness benchmarks based on differences between human-written and LLM-rephrased abstracts. We show that a pooled benchmark may confound pre-existing stylistic variation with AI-generated text, producing substantial distortions across country-field groups even in pre-LLM publications. In contrast, country-field-specific benchmarks attenuate such distortions and provide a more credible baseline for comparison. Applying these methods to publications in 2025 reveals that the pooled benchmark systematically overestimates AI use in certain countries and fields while underestimating it in others. These findings highlight the importance of context-aware measurement for accurate and equitable evaluation of AI use in science.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!