더 높은 정확도, 더 나쁜 추론: 의료 분야 체인 오브 소트 증류 모델의 단계별 성능 분석
Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation
체인 오브 소트(CoT) 증류는 작은 모델을 학습시켜 교사 모델의 추론 과정을 모방하도록 하는 기술이지만, 일반적으로 정확도를 포함한 최종 답변 지표로 평가됩니다. 본 연구에서는 답변 품질이 향상됨에 따라 추론 과정도 개선되는지 질문합니다. 의료 분야 질의응답에서 짧은 답변 옵션으로 인해 임상적 근거가 충분히 명시되지 않는 경우가 많습니다. DeepSeek-V3 계열의 교사 모델로부터 증류된 Qwen3-8B 학생 모델은 MedQA-USMLE 평가 지표(SC@64: 74.7%에서 84.4%; 예상 보정 오류(ECE): 0.096에서 0.034)를 개선했습니다. 그러나 Kimi-K2.6 스타일을 고려하지 않는 LLM 평가 도구를 사용한 분석 결과, 모델이 답변에 대해 '불참'하지 않은 단계에서의 오차율은 30.6%에서 50.3%로 증가했습니다. 이 연구에서는 의료 분야의 특정 사례에서 답변 품질과 추론 과정의 사실성 간에 반대 방향으로 움직이는 현상을 확인했습니다. 이러한 패턴은 평가자, 교사 모델의 강점, 학생 모델의 크기 및 계열, 의료 분야 벤치마크, 그리고 스타일, 세분화, 그리고 정답 여부와 관련된 다양한 제어 변수를 통해 일관적으로 나타났습니다. 임상 전문가가 수행한 150단계의 검토에서도 동일한 결과를 재현했습니다. 추가 분석을 통해 문제의 범위를 좁힌 결과, 간결한 답변이 근거를 충분히 설명하지 못할 때 발생하며, 숙련된 학생 모델이 전문가와 유사한 형식을 모방하더라도 각 단계별 주장을 신뢰성 있게 뒷받침하지 못하는 경우에 문제가 발생하는 것으로 나타났습니다. 기존의 답변 지표 및 집계 불확실성 비율은 이러한 변화를 제대로 반영하지 못합니다. 따라서 체인 오브 소트 추론 과정을 공개하거나 재사용할 때는 답변 수준의 지표만으로는 충분하지 않습니다.
Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics including accuracy. We ask whether gains in answer quality are accompanied by improvements in the trace. In medical QA, where short answer options can leave a richer clinical justification under-specified, a Qwen3-8B student distilled from a DeepSeek-V3-family teacher improves on MedQA-USMLE answer metrics (SC@64 74.7% to 84.4%; expected calibration error (ECE) 0.096 to 0.034). Yet under a Kimi-K2.6 style-blind LLM-judge audit, its error rate over non-abstained steps rises from 30.6% to 50.3%. In this primary medical setting, answer quality and trace factuality move in opposite directions. This before--after pattern persists across evaluators, teacher strengths, student scales and families, medical benchmarks, and style, segmentation, and answer-correctness controls. A 150-step blinded audit by a clinical expert reproduces the same ordering. Boundary checks narrow the scope of the claim: the risk appears when a compact answer under-constrains the rationale and a capable student can imitate expert-like form without reliably grounding each local claim. Standard answer metrics and aggregate hedging rates do not reveal the shift. When such traces are released or reused, answer-level metrics alone are insufficient.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.