2607.05734v1 Jul 07, 2026 cs.IR

SCOReD: 추천 시스템에서의 학생 모델 맞춤형 추론 최적화

SCOReD: Student-Aware CoT Optimization for Recommendation Distillation

Yufei Li
Yufei Li
Citations: 11
h-index: 2
Xiaohan Wei
Xiaohan Wei
Citations: 31
h-index: 2
Chongling Sun
Chongling Sun
Citations: 40
h-index: 2
Sandeep Pandey
Sandeep Pandey
Citations: 1,698
h-index: 21
Luke Simon
Luke Simon
Citations: 24
h-index: 3
Yunchen Pu
Yunchen Pu
Citations: 227
h-index: 1
Fei Tian
Fei Tian
Citations: 282
h-index: 7
F. Shyu
F. Shyu
Citations: 53
h-index: 4
Xi Liu
Xi Liu
Citations: 75
h-index: 3
H. S. Shahgir
H. S. Shahgir
Citations: 103
h-index: 6
Yue Dong
Yue Dong
Citations: 63
h-index: 4

추천 시스템 분야에서 강화 학습(RL)을 위한 전 단계로 활용되는 추론 증류(CoT distillation)는 필수적인 과정이지만, 원본 교사 모델의 답변 기록은 이 작업에 적합하지 않습니다. 거대 교사 모델은 추천 작업을 수행할 때 비정상적으로 높은 불확실성을 보이는데, 정답을 반복적으로 확인하지만 수정하지 않는 경향이 있습니다. 이러한 답변 기록으로 지도 학습을 하면 학생 모델은 장황한 답변을 생성하며 초기 추론을 거의 수정하지 않습니다. 또한, 추천 시스템 분야의 특성상 교사 모델의 추론 기록은 소규모 학생 언어 모델(LLM)에게 매우 이질적인 데이터입니다. 본 논문에서는 추천 시스템에 특화된 추론 최적화 프레임워크인 Student-Aware CoT Optimization for Recommendation Distillation (SCOReD)를 제안합니다. SCOReD는 먼저 각 교사 모델의 답변 기록을 유형별 세그먼트로 분할하고, 학생 LLM의 어텐션 메커니즘을 사용하여 각 세그먼트의 중요도를 평가합니다. 그런 다음, SCOReD는 학생 모델의 출력 길이에 따라 각 세그먼트에 대해 (유지/재작성/병합/삭제) 중 하나의 편집 작업을 동적으로 선택합니다. 따라서 SCOReD는 불필요한 부분을 제거하여 추론 기록의 효율성을 높이는 동시에 정보가 풍부한 부분은 보존하고, 교사 모델의 원본 답변을 학생 모델의 출력 분포에 맞게 조정합니다. SCOReD로 최적화된 추론 데이터를 사용하여 학습하면 학생 모델에게 더 명확한 학습 신호를 제공하며, 기존의 지도 미세 조절(SFT) 방식보다 NDCG가 1.56% 향상되고 Recall@5가 1.9% 향상되는 효과를 보입니다. 또한, SCOReD는 추론 길이를 27.3% 감소시키는 효과도 있습니다.

Original Abstract

Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-suited to this task. Large teachers approach the recommendation task with unusually high reasoning uncertainty, repeatedly rechecking their answers without revising them; supervised fine-tuning on such traces produces verbose students that never revise their initial guess. Furthermore, due to the novelty of the recommendation domain, the teacher's reasoning traces are highly out-of-distribution for the small student LLM. We propose Student-Aware CoT Optimization for Recommendation Distillation (SCOReD), a CoT optimization framework tailored to recommendation that first parses each teacher trace into typed segments and uses the student LLM's attention to score the importance of each segment. Then SCOReD dynamically selects a per-segment edit (KEEP / REWRITE / FUSE / PRUNE) based on the output length and comparative log probability lift of the answer given the edit as per the student. Therefore, SCOReD prunes redundant sections of the reasoning trace while preserving information-dense sections and adapts raw teacher traces to the student's output distribution. Training on SCOReD-optimized CoTs provides a cleaner learning signal to the student model and improves over baseline SFT by 1.56% NDCG and 1.9% Recall@5, while reducing reasoning length by 27.3%.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!