맥락을 고려한 반론은 일반적인 반론보다 더 설득력이 있을 수 있다
Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech
인공지능이 생성하는 반론은 온라인상의 유해 콘텐츠를 줄이고 건설적인 대화를 촉진하는 확장 가능하고 효과적인 전략입니다. 그러나 기존 방식은 획일적인 접근 방식을 채택하여, 대화의 맥락과 대상 사용자의 특징을 간과합니다. 본 연구에서는 다양한 전략을 제안하고 평가하여, 특정 환경에 맞게 조정되고 사용자에게 개인화된 반론을 생성하는 방법을 모색합니다. 구체적으로, 다양한 형태의 맥락 정보와 미세 조정 기술을 통합하는 여러 가지 구성 방식을 탐구합니다. 정량적 지표와 사전에 등록된 혼합 설계 방식의 크라우드소싱 실험을 결합하여 종합적인 평가를 수행했습니다. 견고성을 확보하기 위해 ROUGE, BLEU, BERTScore를 기반으로 한 반론 품질에 대한 알고리즘 측정 방법을 적용했으며, 모든 지표에서 일관된 결과를 관찰했습니다. 또한, 생성된 반론과 대상이 된 유해 메시지의 어떤 특징이 인지되는 설득력에 가장 큰 영향을 미치는지 분석하여, 맥락을 고려한 개입 방식을 어떻게 더욱 효과적으로 만들 수 있는지에 대한 통찰력을 얻었습니다. 연구 결과는 개인화가 효과적일 수 있지만, 항상 그렇지는 않다는 것을 보여줍니다. 대화의 맥락과 사용자 이력을 결합하는 간단한 전략은 인지되는 적절성과 설득력을 향상시키는 반면, 다른 일부 맥락화 전략은 인간이 인지하는 반론 품질을 저하시키는 것으로 나타났습니다. 종합적으로 볼 때, 이러한 결과는 더욱 개인화되고 효과적이며 책임감 있는 반론 시스템 개발을 위한 실질적인 지침을 제공하며, 궁극적으로 온라인 콘텐츠 관리 분야에서 인간과 인공지능의 협력을 발전시키는 데 기여합니다.
AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.