LLM 탐지 기술의 개입 효과: 전략적 사용자 행동 하에서의 파급 영향
LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior
LLM(Large Language Model) 활용이 보편화됨에 따라, LLM이 생성한 콘텐츠를 탐지하는 데 대한 관심이 높아지고 있습니다. 이러한 탐지는 LLM 탐지 도구나 언어 패턴 기반의 휴리스틱을 통해 이루어집니다. 탐지 기술은 단순히 특정 속성을 식별하는 것뿐만 아니라, LLM 사용량 및 출력 품질과 같은 하위 지표에도 영향을 미치는 개입으로 작용합니다. 본 연구에서는 불완전한 LLM 탐지기가 이러한 하위 지표에 예상치 못한 영향을 미치는 것을 보여줍니다. 이는 사용자들이 LLM을 활용하는 방식과 콘텐츠를 후처리하는 방식에 대한 인센티브를 왜곡하기 때문입니다. 우리는 사용자들이 LLM의 활용 정도와, 탐지될 가능성을 줄이기 위해 콘텐츠를 어떻게 수정하는지를 모델링했습니다. 이 모델을 통해, LLM 탐지가 역설적으로 인간 사용자의 LLM 사용량을 증가시킬 수 있음을 보여줍니다. 또한, 탐지되는 속성이 감소하면 출력 품질이 향상되더라도, LLM 탐지 기술 도입은 오히려 사용자들의 낮은 품질의 결과물을 생성하도록 유도할 수 있습니다. 반면, 탐지 기술은 “증가 후 감소”라는 명확한 패턴을 만들어내며, 이는 arXiv 초록의 단어 빈도 분석에서 경험적으로 확인되었습니다. 전반적으로 본 연구는 LLM 탐지가 LLM 사용 및 출력 품질에 어떤 왜곡을 가져올 수 있는지 보여주며, LLM 탐지 기술이 하위 지표에 미치는 개입으로서 작용할 때 발생할 수 있는 문제점을 밝혀냅니다.
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.