CROP: 반사실적 추론을 통한 선택적 온폴리시 증류를 위한 작업 관련성
CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
온폴리시 증류(OPD)는 현재 정책에서 샘플링된 경로를 기반으로 학생 언어 모델을 학습시키지만, 서로 다른 감독 가치를 가진 응답 토큰에 동일한 중요도를 부여합니다. 선택적 OPD는 각 응답 토큰의 추정된 학습 가치에 따라 감독 신호를 비균등하게 할당하여 이러한 한계를 극복합니다. 그러나 대부분의 기존 기준은 불확실성 또는 교사-학생 간 불일치와 같이 최적화 필요성에 주로 초점을 맞추고 있으며, 현재 입력의 의미 내용과 관련된 작업 관련성은 보완적인 차원으로 직접적으로 명시되지 않았습니다. 이러한 격차를 해소하기 위해, 우리는 반사실적 추론을 통한 온폴리시 증류(CROP)라는 새로운 방법을 제안합니다. CROP은 각 소스 프롬프트에 대해 검증된 원본-패러프레이즈-반사실적 삼중항을 구성하고, 학생 모델의 출력을 고정시킨 상태에서, 각 응답 위치가 작업 관련 조건 변화에 대한 민감도와 의미를 보존하는 재작성에 대한 민감도를 기준으로 측정합니다. 매칭된 실험 결과는 CROP이 무작위 선택 또는 가장 낮은 관련성을 가진 선택보다 더 유용한 감독 신호를 식별한다는 것을 보여줍니다. 또한, 구성 요소 비교 분석은 반사실적 민감도와 패러프레이즈 교정 모두의 가치를 확인했습니다. 두 가지 교사-학생 설정에서, CROP은 가장 강력한 비-CROP 선택 방법보다 각각 1.92점과 2.96점을 향상된 성능을 보였습니다. 이러한 결과는 작업 관련성이 선택적 OPD를 위한 보완적인 기준임을 뒷받침하며, CROP이 토큰 수준의 감독 신호를 할당하는 모델 내부적인, 컨트라스트 특유의 방법임을 입증합니다.
On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation by allocating supervision non-uniformly across response tokens according to their estimated training value. Most existing criteria, however, focus primarily on optimization need, such as uncertainty or teacher-student disagreement, while task relevance, namely whether the supervision is tied to the semantic content of the current input, remains less directly characterized as a complementary dimension. To address this gap, we introduce Counterfactual Relevance for On-Policy Distillation (CROP), which operationalizes task relevance through a paraphrase-calibrated counterfactual sensitivity margin. For each source prompt, CROP constructs a validated original-paraphrase-counterfactual triplet, holds the student rollout fixed, and measures each response position by its sensitivity to a task-relevant condition change calibrated by its sensitivity to a meaning-preserving rewrite. Matched selection controls show that CROP identifies more useful supervision positions than random or lowest-relevance selection, while component comparisons confirm the value of both counterfactual sensitivity and paraphrase calibration. Across two teacher-student settings, CROP improves aggregate performance by 1.92 and 2.96 points over the strongest non-CROP selector. These results support task relevance as a complementary criterion for selective OPD and establish CROP as a model-internal, contrast-specific method for allocating token-level supervision.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.