AI 어시스턴트의 과도한 도움: 문제점과 해결 방안
AI Assistants Overassist
최근 대규모 언어 모델(LLM)은 사용자들이 문제를 해결하는 과정을 지원하고 사고를 돕는 튜터 및 조력자로서 활용되고 있습니다. AI 어시스턴트의 안내는 학습을 촉진하고 사고 능력을 향상시키는 데 도움이 될 수 있지만, 그 효과는 AI가 어떻게 도움을 제공하느냐에 따라 달라집니다. 예를 들어, 너무 이르거나 빈번한 개입은 실제 학습과 인지적 참여를 저해할 수 있습니다. 하지만 AI 시스템이 문제 해결 과정에서 개입 시점을 어떻게 결정하는지에 대한 이해는 아직 부족합니다. 본 연구에서는 LLM의 개입 효과를 평가하기 위한 시뮬레이션 기반 벤치마크인 Int-Bench를 소개합니다. Int-Bench는 '학생'이 문제를 풀고, '교사'가 학생의 사고 과정을 모니터링하며, 언제, 어떻게 개입할지 결정하는 상황을 시뮬레이션합니다. 본 연구에서는 코드 디버깅, 수학, 퍼즐 문제 해결이라는 세 가지 영역에서 LLM 교사의 개입 빈도와 시점, 그리고 즉각적인 과제 성공 및 새로운 문제에 대한 일반화 능력에 미치는 영향을 평가했습니다. 또한 LLM과 인간의 개입 방식을 비교한 결과, LLM은 인간보다 더 자주, 더 일찍 개입하는 경향이 있었습니다. 더욱이, 인간과는 달리 LLM은 종종 구체적인 힌트 대신 완전한 해결책을 제공합니다. 이러한 결과는 현재 LLM 어시스턴트가 깊이 있는 학습과 장기적인 성공에 필요한 사고 과정을 지원하기보다는 단기적인 성공을 최적화하는 경향이 있음을 시사합니다.
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assistants can scaffold thinking and foster learning, such benefits depend on how they help--for instance, intervening too early or too frequently may hinder true learning and cognitive engagement. Yet how AI systems navigate intervention decisions during problem-solving remains poorly understood. Here, we introduce Int-Bench, a simulation-based benchmark for evaluating LLM interventions during learning. Int-Bench simulates a "student" solving a problem while a "teacher" monitors the student's reasoning and decides whether, when, and how to intervene. Across three domains--code debugging, mathematics, and brain teasers--we evaluate LLM teachers on the frequency and timing of interventions, as well as their impact on both immediate task success and generalization to new problems. We also compare LLMs to humans, finding that LLMs intervene more frequently and earlier than humans. Moreover, in contrast to humans, they tend to provide complete solutions rather than targeted hints. These findings suggest that current LLM assistants often optimize for short-term success rather than supporting the reasoning processes needed for deeper learning and long-term success.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.