2608.03859v1 Aug 04, 2026 cs.CL

표현 유사성 너머: 생성적 표절 탐지를 위한 소스 조건부 설명 길이 기반 이득 및 후보 소스 재순위화

Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Wenxuan Xie
Wenxuan Xie
Citations: 74
h-index: 4
Pei-Fu Guo
Pei-Fu Guo
Citations: 74
h-index: 2

대규모 언어 모델(LLM)은 학문 윤리와 동료 평가에 어려움을 야기합니다. 그러나 생성적 표절 탐지는 아직 연구가 부족하고 해결되지 않은 과제로 남아 있습니다. 기존의 LLM-생성 텍스트 탐지 연구는 AI의 개입 여부를 판단하는 데 초점을 맞추고 있으며, 이는 허용될 수 있는 반면, 유사성을 기반으로 하는 방법은 광범위한 재작성과 다중 소스 통합 이후에는 효과가 떨어집니다. 확률적 예측의 설명 길이 관점에서, 관련 부가 정보는 대상 시퀀스의 코드 길이를 줄일 수 있다는 점에 착안하여, 우리는 소스 조건부 설명 길이 기반 이득(Source-Conditioned Description-Length Gain, SCDG)이라는 새로운 프레임워크를 제안합니다. SCDG는 훈련 과정 없이, 검증된 언어 모델이 의심스러운 문서 $P$의 설명 길이를 계산하는데, 이때 후보 소스 $S$가 존재하는 경우와 존재하지 않는 경우를 비교합니다. 이러한 비교를 통해 토큰 수준의 로그-우도 이득을 얻을 수 있으며, 이는 $S$가 제공하는 예측 증거의 증가량을 측정합니다. 우리는 생성적 표절 탐지를 위한 PAN at CLEF 벤치마크에서 SCDG를 평가했습니다. PAN 2025 데이터를 기반으로 한 쌍별 벤치마크에서 SCDG는 0.92의 정밀도, 0.97의 재현율 및 0.94의 F1 값을 달성하여 모든 기준 모델을 능가했습니다. 또한, PAN 2026의 다중 소스 검색 작업에서는 0.83의 nDCG@10과 0.96의 Recall@100을 기록하며 모든 기준 모델을 뛰어넘었습니다. 동일한 주제와 사건에 대한 Multi-News 테스트에서, 보정된 이득 분포를 사용하는 SCDG 분류기는 전체 쌍 중 0.125%에 대해서만 소스 재사용을 예측하여, 이러한 평가 프로토콜 하에서의 주제적 중복성에 대한 강건성을 입증합니다. 이러한 결과는 SCDG가 광범위한 변환 과정에서도 소스별 콘텐츠 재사용을 감지하는 데 유용하고 토큰 단위로 분석 가능한 신호임을 보여줍니다.

Original Abstract

Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets AI involvement, which may be permissible, rather than source reuse, while similarity-based methods struggle after extensive rewriting and multi-source synthesis. Motivated by the description-length view of probabilistic prediction, in which relevant side information can reduce a target sequence's code length, we introduce Source-Conditioned Description-Length Gain (SCDG), a directional, training-free framework that contrasts a frozen language model's description length of a suspicious document $P$ with and without a candidate source $S$. This contrast yields token-level log-likelihood gains that measure the incremental predictive evidence supplied by $S$. We evaluate SCDG on the PAN at CLEF benchmarks for generative plagiarism. On a PAN 2025-derived pairwise benchmark, SCDG achieves 0.92 Precision, 0.97 Recall, and 0.94 F1, outperforming all baselines; on PAN 2026's multi-source retrieval task, it reaches 0.83 nDCG@10 and 0.96 Recall@100, surpassing all baselines. On a same-topic, same-event Multi-News test, the calibrated gain-distribution SCDG classifier predicts source reuse for only $0.125\%$ of pairs, supporting robustness to topical overlap under this evaluation protocol. These results establish SCDG as a unified and token-decomposable signal for source-specific content reuse under extensive transformation.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!