2607.04572v1 Jul 06, 2026 cs.AI

LLM 기반 교육 튜터에서 체인 오브 소트(Chain-of-Thought) 감사 기법을 활용한 정답 중심 추론 탐지

Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing

Bowen Liu
Bowen Liu
Citations: 4
h-index: 1
Bonan Shen
Bonan Shen
Citations: 8
h-index: 1
Dingyan Shang
Dingyan Shang
Citations: 0
h-index: 0
Tao Ning
Tao Ning
Citations: 6
h-index: 2
You Wang
You Wang
Citations: 52
h-index: 2

대규모 언어 모델(LLM) 기반 튜터는 종종 유창하고 단계별 설명을 생성하지만, 올바르고 교육적으로 구성된 응답이 학생에게 제시된 문제로부터 도출되었다는 것을 보장하지 않습니다. 실제 튜터링 시스템에서 모델은 교사 노트, 정답 키, 채점 기준 또는 검색된 솔루션 자료에 접근할 수 있습니다. 본 연구에서는 이러한 개인 정보가 튜터의 설명을 정답 중심으로 만들 수 있는지 조사합니다. 즉, 서면 설명이 그 내용을 뒷받침하기 전에 최종 정답이 행동적으로 사용 가능한지 여부를 확인합니다. 체인 오브 소트 전방향(prefix)의 유효성을 검증하는 'Truncated Reasoning AUC Evaluation (TRACE)' 기법을 사용하여 1000개의 GSM8K 테스트 문제를 세 가지 조건으로 비교했습니다: 문제만 제시, 올바른 정답 키 제시, 잘못된 정답 키 제시. 생성된 설명의 고정된 비율에서 모델이 즉시 답하도록 강제하고 응답을 표준 정수 답변과 비교합니다. Qwen2.5-3B-Instruct 모델에서 정답 키 접근은 TRACE AUC 중앙값을 0.375에서 0.900으로 높이고, 1000개의 사례 중 997건에서 첫 번째 10% 전방향에서 표준 정답이 나타나는 것을 확인했습니다. 문제만 제시된 경우와 정답 키를 사용한 경우 모두 올바른 답변으로 끝나는 746개 예시에서도 이러한 효과는 지속되었습니다. 본 연구 결과는 체인 오브 소트 감사 기법이 수학 튜터링 설명에서 정답 중심 추론을 진단하는 데 유용한 경량화된 방법임을 뒷받침합니다.

Original Abstract

Large language model (LLM) tutors often produce fluent step-by-step explanations, but a correct and pedagogically formatted response does not guarantee that the answer was derived from the student-facing problem. In realistic tutoring systems, the model may also have access to teacher notes, answer keys, rubrics, or retrieved solution artifacts. We study whether such private answer information can make tutor explanations answer-driven: the final answer is behaviorally available before the written explanation has justified it. Using Truncated Reasoning AUC Evaluation (TRACE), which probes how early a chain-of-thought prefix can pass a verifier, we evaluate 1000 GSM8K test problems under three paired tutoring contexts: question-only, correct answer-key, and wrong answer-key. At fixed fractions of each generated explanation, we force the model to answer immediately and verify the response against the gold numeric answer. With Qwen2.5-3B-Instruct, answer-key access raises median TRACE AUC from 0.375 to 0.900 and makes the gold answer available at the first 10% prefix in 997 of 1000 cases. The effect remains strong on the 746 examples where both question-only and answer-key explanations end with the correct answer. These results support truncated CoT auditing as a lightweight process-level diagnostic for answer-driven reasoning in math tutoring explanations.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!