2607.29124v1 Jul 31, 2026 cs.CV

SciFigPlag-Bench: 과학적 그림 표절 탐지를 위한 출처 추적 기능 평가 기준

SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection

Minghao Yang
Minghao Yang
Citations: 73
h-index: 4
Zhiying Cui
Zhiying Cui
Citations: 46
h-index: 4
Linlin Gao
Linlin Gao
Citations: 89
h-index: 3
Pengyuan Li
Pengyuan Li
Citations: 19
h-index: 2
Jie Liu
Jie Liu
Citations: 0
h-index: 0

과학적 그림은 종종 과학적 결과에 대한 시각적 증거를 담고 있지만, 그림 표절 문제는 아직 체계적인 다중 모드 평가 문제로서 제대로 연구되지 않았습니다. 본 논문에서는 SciFigPlag-Bench라는, 학술 문서 내의 과학적 그림에 대한 출처 추적 기능을 평가하는 새로운 기준을 제시합니다. 일반적인 이미지 유사성 또는 이미지 포렌식 벤치마크와 달리, SciFigPlag-Bench는 의심스러운 그림이 특정 원본 그림에서 어떤 증거를 재사용했는지, 재사용된 내용이 어떻게 변형되었는지, 그리고 재사용된 증거가 어디에 나타나는지를 평가합니다. 우리는 재사용되는 요소와 그 변환 방식을 분리하는 계층적 분류 체계를 도입하여, 전체 그림 또는 부분 그림의 재사용과 같은 물질 보존적 재사용뿐만 아니라 데이터 재표현 및 구조적 재구성과 같은 추상적인 내용 재사용까지 포함합니다. 이 분류 체계에 따라, 우리는 2,582개의 양성 쌍과 2,541개의 음성 쌍으로 구성된 하이브리드 벤치마크를 구축했으며, 여기에는 문서화된 실제 사례, 분류 체계를 따른 합성 예제 및 시각적으로 유사한 부정 예제가 포함됩니다. 이 벤치마크는 쌍별 탐지, 출처 추적, 계층적 재사용 유형 분류 및 재사용 대응 관계 위치 파악의 네 가지 진단 작업을 지원합니다. 다양한 시각-언어 모델을 사용한 실험은 초기 성능 기준을 확립하고, 미세 수준의 출처 추론, 재사용 유형 이해 및 공간적 증거 연결과 관련된 지속적인 과제를 드러냅니다.

Original Abstract

Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal evaluation problem. We present SciFigPlag-Bench, a benchmark for provenance-aware reasoning over scientific figures in scholarly documents. Unlike general image-similarity or image-forensics benchmarks, SciFigPlag-Bench evaluates whether a suspicious figure reuses evidence from a specific source figure, how the reused content has been transformed, and where the reused evidence appears. We introduce a factorized taxonomy that separates what is reused from how it is transformed, covering material-preserving reuse, such as full-figure and subfigure reuse, as well as abstract-content reuse, such as data re-expression and structural redraw. Guided by this taxonomy, we construct a hybrid benchmark with 2,582 positive pairs and 2,541 negative pairs, combining documented real-world cases, taxonomy-guided synthetic examples, and visually similar negatives. The benchmark supports four diagnostic tasks: pairwise detection, source attribution, hierarchical reuse-type classification, and reuse correspondence localization. Experiments with diverse vision-language models establish initial baselines and reveal persistent challenges in fine-grained provenance reasoning, reuse-type understanding, and spatial evidence grounding.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!