2602.24080v1 Feb 27, 2026 cs.AI

인간인가, 기계인가? 음성-음성 상호작용을 위한 초기 튜링 테스트

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

Xiang Li
Xiang Li
Citations: 26
h-index: 2
Jiabao Gao
Jiabao Gao
Citations: 61
h-index: 3
Sipei Lin
Sipei Lin
Citations: 1
h-index: 1
Xu Zhou
Xu Zhou
Citations: 2
h-index: 1
Chi Zhang
Chi Zhang
Citations: 90
h-index: 5
Bo Cheng
Bo Cheng
Citations: 36
h-index: 4
Jiale Han
Jiale Han
Hong Kong University of Science and Technology
Citations: 380
h-index: 9
Benyou Wang
Benyou Wang
Citations: 446
h-index: 2

인간과 유사한 대화형 에이전트 개발은 오랫동안 튜링 테스트에 의해 이끌어져 왔습니다. 현대 음성-음성(S2S) 시스템의 경우, 인간처럼 대화할 수 있는지 여부는 중요한 질문이지만 아직 명확하게 답변되지 않았습니다. 이를 해결하기 위해, 우리는 S2S 시스템을 위한 최초의 튜링 테스트를 수행하고, 9개의 최첨단 S2S 시스템과 28명의 인간 참가자 간의 대화에 대한 2,968건의 인간 평가를 수집했습니다. 우리의 결과는 명확한 결론을 제시합니다. 즉, 현재 평가된 S2S 시스템 중 어느 것도 튜링 테스트를 통과하지 못하며, 이는 인간과 유사성 측면에서 상당한 격차가 있음을 보여줍니다. 이러한 실패의 원인을 분석하기 위해, 우리는 18가지의 세분화된 인간 유사성 차원을 개발하고, 수집된 대화에 대한 다중 평가를 수행했습니다. 분석 결과, 주요 문제는 의미 이해가 아니라 비언어적 특징, 감정 표현, 대화 페르소나에서 비롯되는 것으로 나타났습니다. 또한, 기존의 AI 모델이 튜링 테스트 판별자로 사용될 때 신뢰성이 낮다는 것을 확인했습니다. 이에 대한 대응으로, 우리는 세분화된 인간 유사성 평가를 활용하여 정확하고 투명한 인간-기계 구분을 제공하는 해석 가능한 모델을 제안합니다. 이 모델은 자동화된 인간 유사성 평가를 위한 강력한 도구를 제공합니다. 본 연구는 S2S 시스템을 위한 최초의 인간 유사성 평가를 확립하고, 단순한 성공/실패 결과를 넘어 상세한 진단 정보를 제공함으로써, 대화형 AI 시스템의 인간과 유사한 발전을 위한 기반을 마련합니다.

Original Abstract

The pursuit of human-like conversational agents has long been guided by the Turing test. For modern speech-to-speech (S2S) systems, a critical yet unanswered question is whether they can converse like humans. To tackle this, we conduct the first Turing test for S2S systems, collecting 2,968 human judgments on dialogues between 9 state-of-the-art S2S systems and 28 human participants. Our results deliver a clear finding: no existing evaluated S2S system passes the test, revealing a significant gap in human-likeness. To diagnose this failure, we develop a fine-grained taxonomy of 18 human-likeness dimensions and crowd-annotate our collected dialogues accordingly. Our analysis shows that the bottleneck is not semantic understanding but stems from paralinguistic features, emotional expressivity, and conversational persona. Furthermore, we find that off-the-shelf AI models perform unreliably as Turing test judges. In response, we propose an interpretable model that leverages the fine-grained human-likeness ratings and delivers accurate and transparent human-vs-machine discrimination, offering a powerful tool for automatic human-likeness evaluation. Our work establishes the first human-likeness evaluation for S2S systems and moves beyond binary outcomes to enable detailed diagnostic insights, paving the way for human-like improvements in conversational AI systems.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!