채팅 기반 지원만으로는 충분하지 않을 수 있다: 수학 증명 학습을 위한 대화형 및 내장형 LLM 피드백 비교
Chat-Based Support Alone May Not Be Enough: Comparing Conversational and Embedded LLM Feedback for Mathematical Proof Learning
본 연구는 학부 이산수학 과정을 위한 LLM 기반 튜터링 시스템인 GPTutor를 평가한다. 이 시스템은 학생들이 작성한 증명 시도에 대해 내장형 피드백을 제공하는 구조화된 증명 검토 도구와 수학 질문을 위한 챗봇이라는 두 가지 LLM 지원 도구를 통합한다. 148명의 학생을 대상으로 한 시차 접근(staggered-access) 연구에서, 실험군만 시스템을 사용할 수 있었던 기간 동안 조기 접근은 더 높은 과제 성취도와 연관이 있었으나, 이러한 성취도 향상이 시험 점수로 전이되는 것은 관찰하지 못했다. 사용 로그에 따르면 자기 효능감과 사전 시험 성적이 낮은 학생일수록 두 구성 요소를 더 자주 사용한 것으로 나타났다. 인간 코딩을 통해 생성되고 자동 분류기를 사용하여 확장된 세션 수준의 행동 레이블은 학생들이 챗봇과 상호작용하는 방식(예: 정답 탐색 또는 도움 탐색)을 특징짓는다. 사전 성적과 자기 효능감을 통제한 모델에서 높은 챗봇 사용률과 정답 탐색 행동은 이후 중간고사 성적과 부정적인 연관을 보인 반면, 증명 검토 도구의 사용은 감지할 만한 독립적인 연관성을 나타내지 않았다. 이러한 결과는 종합적으로 채팅 기반 지원만으로는 수학 증명 학습 결과에 대한 독립적인 평가로의 전이를 안정적으로 지원하지 못할 수 있는 반면, 작업에 기반한 구조화된 피드백은 학습 저하와의 연관성이 더 적어 보인다는 점을 시사한다.
We evaluate GPTutor, an LLM-powered tutoring system for an undergraduate discrete mathematics course. It integrates two LLM-supported tools: a structured proof-review tool that provides embedded feedback on students' written proof attempts, and a chatbot for math questions. In a staggered-access study with 148 students, earlier access was associated with higher homework performance during the interval when only the experimental group could use the system, while we did not observe this performance increase transfer to exam scores. Usage logs show that students with lower self-efficacy and prior exam performance used both components more frequently. Session-level behavioral labels, produced by human coding and scaled using an automated classifier, characterize how students engaged with the chatbot (e.g., answer-seeking or help-seeking). In models controlling for prior performance and self-efficacy, higher chatbot usage and answer-seeking behavior were negatively associated with subsequent midterm performance, whereas proof-review usage showed no detectable independent association. Together, the findings suggest that chatbot-based support alone may not reliably support transfer to independent assessment of math proof-learning outcomes, whereas work-anchored, structured feedback appears less associated with reduced learning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.