Triospect: 다양한 공격에 대한 강력한 통계 기반 AI 생성 텍스트 탐지 시스템을 위한 3차원 프레임워크
Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks
기존의 AI 생성 텍스트 탐지기는 텍스트 특성을 조작하는 공격에 취약합니다. 본 연구에서는 주어진 텍스트 내에서 내용(핵심 아이디어)과 표현(스타일 요소)이라는 추가적인 관점을 활용하여 새로운 Triospect 탐지 프레임워크를 제안합니다. 17개의 공격, 12개의 도메인, 그리고 17개의 소스 모델을 포함하는 두 가지 벤치마크 실험 결과, Triospect는 이러한 공격에 대해 강력한 성능을 보이는 것을 확인했습니다. 특히 Humanize-16K 데이터 세트의 공격 이후 부분에서 기존 최적 성능보다 각각 22.3% (AUROC) 및 13% (TPR01) 향상되었으며, adversarial RAID 데이터 세트에서는 각각 9.1% (AUROC) 및 22% (TPR01) 향상을 보였습니다. 본 프레임워크는 공격에 대한 탐지 신뢰성을 향상시키는 통계적 방법론 분야에서 중요한 진전을 나타냅니다. 저희가 생성한 데이터와 코드는 https://github.com/baoguangsheng/triospect 에서 확인할 수 있습니다.
Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics. In this study, we propose a novel Triospect Detection Framework by using additional perspectives of content (core ideas) and expression (stylistic elements) within a given text. Experiments on two benchmarks involving 17 attacks, 12 domains, and 17 source models demonstrate that Triospect is robust against these attacks. It improves the strong baseline by a significant margin of 22.3% (AUROC) and 13% (TPR01) on the Humanize-16K after-attack subset, and by 9.1% (AUROC) and 22% (TPR01) on the adversarial RAID. This framework marks a pioneering effort in statistical methods to enhance detection reliability against attacks. We release our data and code at https://github.com/baoguangsheng/triospect.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.