2602.18807v1 Feb 21, 2026 cs.HC

채팅 기반 지원만으로는 충분하지 않을 수 있다: 수학 증명 학습을 위한 대화형 및 내장형 LLM 피드백 비교

Chat-Based Support Alone May Not Be Enough: Comparing Conversational and Embedded LLM Feedback for Mathematical Proof Learning

Eason Chen
Eason Chen
Citations: 144
h-index: 7
Sophia Judicke
Sophia Judicke
Citations: 2
h-index: 1
Kayla Beigh
Kayla Beigh
Citations: 2
h-index: 1
Xinyi Tang
Xinyi Tang
Citations: 30
h-index: 4
I-Ting Wang
I-Ting Wang
Citations: 7
h-index: 1
Nina Yuan
Nina Yuan
Citations: 0
h-index: 0
Zimo Xiao
Zimo Xiao
Citations: 9
h-index: 2
Chuangji Li
Chuangji Li
Citations: 8
h-index: 2
Shizhuo Li
Shizhuo Li
Citations: 8
h-index: 2
Reed Luttmer
Reed Luttmer
Citations: 2
h-index: 1
Shreya Singh
Shreya Singh
Citations: 10
h-index: 2
M. Yampolsky
M. Yampolsky
Citations: 8
h-index: 2
N. Parikh
N. Parikh
Citations: 10
h-index: 2
Yvonne Zhao
Yvonne Zhao
Citations: 23
h-index: 2
M. Chen
M. Chen
Citations: 32
h-index: 1
Scarlett Huang
Scarlett Huang
Citations: 8
h-index: 2
Anishka Mohanty
Anishka Mohanty
Citations: 9
h-index: 2
John R Mackey
John R Mackey
Citations: 1
h-index: 1
Kenneth R. Koedinger
Kenneth R. Koedinger
Citations: 80
h-index: 5
Gregory K. Johnson
Gregory K. Johnson
Citations: 18
h-index: 2
Jionghao Lin
Jionghao Lin
Citations: 885
h-index: 13

본 연구는 학부 이산수학 과정을 위한 LLM 기반 튜터링 시스템인 GPTutor를 평가한다. 이 시스템은 학생들이 작성한 증명 시도에 대해 내장형 피드백을 제공하는 구조화된 증명 검토 도구와 수학 질문을 위한 챗봇이라는 두 가지 LLM 지원 도구를 통합한다. 148명의 학생을 대상으로 한 시차 접근(staggered-access) 연구에서, 실험군만 시스템을 사용할 수 있었던 기간 동안 조기 접근은 더 높은 과제 성취도와 연관이 있었으나, 이러한 성취도 향상이 시험 점수로 전이되는 것은 관찰하지 못했다. 사용 로그에 따르면 자기 효능감과 사전 시험 성적이 낮은 학생일수록 두 구성 요소를 더 자주 사용한 것으로 나타났다. 인간 코딩을 통해 생성되고 자동 분류기를 사용하여 확장된 세션 수준의 행동 레이블은 학생들이 챗봇과 상호작용하는 방식(예: 정답 탐색 또는 도움 탐색)을 특징짓는다. 사전 성적과 자기 효능감을 통제한 모델에서 높은 챗봇 사용률과 정답 탐색 행동은 이후 중간고사 성적과 부정적인 연관을 보인 반면, 증명 검토 도구의 사용은 감지할 만한 독립적인 연관성을 나타내지 않았다. 이러한 결과는 종합적으로 채팅 기반 지원만으로는 수학 증명 학습 결과에 대한 독립적인 평가로의 전이를 안정적으로 지원하지 못할 수 있는 반면, 작업에 기반한 구조화된 피드백은 학습 저하와의 연관성이 더 적어 보인다는 점을 시사한다.

Original Abstract

We evaluate GPTutor, an LLM-powered tutoring system for an undergraduate discrete mathematics course. It integrates two LLM-supported tools: a structured proof-review tool that provides embedded feedback on students' written proof attempts, and a chatbot for math questions. In a staggered-access study with 148 students, earlier access was associated with higher homework performance during the interval when only the experimental group could use the system, while we did not observe this performance increase transfer to exam scores. Usage logs show that students with lower self-efficacy and prior exam performance used both components more frequently. Session-level behavioral labels, produced by human coding and scaled using an automated classifier, characterize how students engaged with the chatbot (e.g., answer-seeking or help-seeking). In models controlling for prior performance and self-efficacy, higher chatbot usage and answer-seeking behavior were negatively associated with subsequent midterm performance, whereas proof-review usage showed no detectable independent association. Together, the findings suggest that chatbot-based support alone may not reliably support transfer to independent assessment of math proof-learning outcomes, whereas work-anchored, structured feedback appears less associated with reduced learning.

3 Citations
0 Influential
6.5 Altmetric
35.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!