AI, 운전을 맡겨주세요: 인간과 컴퓨터의 협력적인 질문 답변 시스템에서 위임(delegation) 및 신뢰에 영향을 미치는 요인은 무엇인가?
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
인공지능 시스템은 오류를 범할 수 있으며, 인간 또한 자신의 판단보다 AI를 신뢰할지 여부를 결정하는 과정에서 실수를 할 수 있습니다. 따라서 인간-AI 협력을 향상시키기 위해서는 인간이 언제, 왜, 그리고 어떻게 AI에 의존하는지를 이해해야 합니다. 본 연구에서는 두 가지 뚜렷한 의존 결정 방식을 분석합니다. 첫째는 '위임'으로, AI의 결과를 미리 알지 않고 AI가 자율적으로 작동하도록 하는 시점을 결정하는 것입니다. 둘째는 '채택'으로, AI의 제안을 평가하고 이를 어떻게 활용할지 결정하는 것입니다. 이러한 독립적인 의존 패턴은 협력에 영향을 미치지만, 기존 연구에서는 현실적인 환경에서 동일한 사용자를 대상으로 이 두 가지를 함께 연구하는 경우가 드뭅니다. 본 연구는 인간-AI 팀이 질문 답변 게임에서 경쟁하면서, 인간이 AI 에이전트와 언제, 어떻게 협력할지 선택할 수 있는 상황을 통해 이러한 간극을 해소하고자 합니다. 총 24경기의 경기에서 23명의 숙련된 인간 사용자와 16개의 AI 에이전트를 페어링하여, 387건의 위임 결정과 1440건의 채택 결정을 분석했습니다. 결과적으로, 인간-AI 협력은 AI 또는 인간 단독보다 더 나은 성능을 보이지만, 인간은 최적의 협력 결정을 내리지 못하는 경우가 있었습니다. 즉, 올바른 AI 제안에 충분히 의존하지 않는 경우 (3.9%의 기회 손실)와 AI가 잘못된 정보를 제공할 때 지나치게 의존하는 경우 (1.7%) 모두 발생했습니다. 두 주체 모두 오답을 내놓으며, 인간과 AI가 의견이 다를 때 모델의 신뢰도 추정치는 무작위 수준에 가깝지만, AI 제안이 인간의 초기 부정확한 답변과 일치할 때는 확증 편향으로 인해 의존도가 더욱 낮아지는 경향 (64.5%)을 보였습니다. 이러한 문제를 해결하기 위해, 우리는 신뢰도 교정, 근거 기반 설명 제공, 그리고 사용자가 신뢰를 조정하는 데 도움이 되는 메커니즘을 제안합니다.
AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration requires understanding when, why, and how humans decide to rely on AI. We study two distinct reliance decisions: the delegation choice -- deciding when to let AI act autonomously without knowing its output, and the adoption choice -- evaluating AI suggestions and deciding how to use them. Both of these decoupled reliance patterns shape collaboration, but prior work rarely studies them together in realistic settings with the same users. We address this gap by studying collaborative human--AI teams competing in a question-answering game in which humans can choose when and how to work with AI agents to win. Our 24 matches pair 23 expert humans with 16 AI agents, capturing 387 delegation and 1440 adoption decisions. While human--AI collaboration performs better than either AI or humans alone, humans make suboptimal collaboration decisions, both under-relying on correct AI suggestions (3.9% of opportunities missed) and over-relying when AI misleads them (1.7%). Both parties contribute wrong answers: reported model confidence is near chance when humans and AI disagree, while confirmation bias drives higher under-reliance (64.5%) when an AI suggestion agrees with humans' initial incorrect answer. To close this gap, we recommend calibrated confidence, evidence-grounded explanations, and mechanisms that help users refine trust.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.