2607.19088v1 Jul 21, 2026 cs.CL

DAIS: 의존성 기반 중간 질의응답 감독 학습을 통한 복잡한 추론

DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning

Yu Wang
Yu Wang
Citations: 451
h-index: 11
Zhihu Wang
Zhihu Wang
Citations: 55
h-index: 3
Xicheng Zhang
Xicheng Zhang
Citations: 29
h-index: 4
Ming Fan
Ming Fan
Citations: 20
h-index: 2
Zhiyong Li
Zhiyong Li
Citations: 0
h-index: 0
Caiyue Xu
Caiyue Xu
Citations: 0
h-index: 0
Dahai Hu
Dahai Hu
Citations: 0
h-index: 0
Ting Liu
Ting Liu
Citations: 55
h-index: 3

체인 오브 소트 (Chain-of-thought, CoT) 방식은 중간 추론 과정을 보여주지만, 일반적으로 평탄화된 추론 목표는 단일 추론 순서를 최적화하고, 로컬 결론이 후속 결정에 어떻게 기여해야 하는지에 대한 제한적인 감독을 제공합니다. 본 논문에서는 의존성 기반 중간 질의응답 감독 학습 (Dependency-Aware Intermediate QA Supervision, DAIS)이라는 훈련 시간 프레임워크를 소개합니다. DAIS는 필터링된 교사 추론 과정을 단계별 질의응답 레코드로 변환합니다. 각 중간 레코드에서 이전 결정에 필요한 이전 상태를 기반으로 로컬 답변을 예측하고, 최종 답변 레코드는 원래 작업 형식을 유지하며, 따라서 평가는 원래 입력과 선택적 컨텍스트만 사용합니다. GDPR, AIACT, MedQA 및 FOLIO 데이터셋에서 다양한 Qwen 모델을 사용하여 실험한 결과, DAIS는 answer-only, flat chain-of-thought, 그리고 독립적인 질의응답 기반 모델보다 평균 최종 답변 정확도를 향상시켰습니다. 정책 준수 벤치마크에서는 가장 강력한 DAIS가 아닌 기준 모델 대비 최대 5.6%의 성능 향상을 보았으며, 평균적으로는 4.2%의 성능 향상을 달성했습니다. 통제된 분석 결과, 이전 상태에 대한 올바른 조건부 설정은 단순히 더 긴 목표 또는 추가적인 중간 텍스트를 제공하는 것 이상으로 기여하며, 이는 의존성을 고려한 중간 질의응답을 표준 최종 답변 추론을 위한 경량 보조 감독 신호로 사용될 수 있음을 보여줍니다.

Original Abstract

Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We introduce Dependency-Aware Intermediate QA Supervision (DAIS), a training-time framework that converts filtered teacher rationales into stage-level QA records. Each intermediate record predicts a local answer conditioned on the previous states needed for that decision, while the final-answer record keeps the original task format; evaluation therefore uses only the original input and optional context. Across GDPR, AIACT, MedQA, and FOLIO with multiple Qwen backbones, DAIS improves average final-answer accuracy over answer-only, flat chain-of-thought, and independent-QA baselines. On policy-compliance benchmarks, it achieves a largest gain of 5.6% and an average gain of 4.2% over the strongest non-DAIS baseline. Controlled ablations show that valid previous-state conditioning contributes beyond longer targets or additional intermediate text, supporting dependency-conditioned intermediate QA as a lightweight auxiliary supervision signal for standard final-answer inference.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!