LongCrafter: 증거 그래프 기반 지시문 생성 학습을 통한 다양한 장문 맥락 이해 연구
LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
장문 맥락 이해 능력을 향상시키는 효과적인 방법은 대규모 언어 모델(LLM)에 대한 장문 맥락 지도 학습 데이터(SFT)를 효율적으로 생성하는 것입니다. 그러나 기존 접근 방식은 다음과 같은 세 가지 한계를 가지고 있습니다: 제한적인 작업 범위, 불충분한 지시문의 난이도, 그리고 신뢰성 있는 감독 부족입니다. 본 논문에서는 계층적 작업 분류 체계와 증거 기반 파이프라인을 결합한 구조화된 생성 프레임워크인 extbf{LongCrafter}를 제안합니다. LongCrafter는 장문 맥락 이해를 지역/피상적인 수준과 전체/심오적인 수준으로 나누어 32가지 세분화된 작업 유형을 정의하며, 이는 전역적 생성 우선순위를 제공합니다. 이 분류 체계를 기반으로 LongCrafter는 작업에 맞춰 조정된 장문 맥락을 구성하고, 단락 간의 의존성을 모델링하는 명시적인 증거 그래프로 이를 분해합니다. 또한, 위치한 증거 범위를 엄격하게 참조하여 지시문-응답 쌍을 생성함으로써, 제어 가능한 난이도와 신뢰성 있는 추론을 보장합니다. LongCrafter 데이터로 미세 조정된 모델은 Qwen2.5-7B 및 LLaMA-3.1-8B 모델에서 LongBench, LongBench~v2, 그리고 LooGLE 벤치마크에서 기존 SFT 기준 모델과 공식적으로 후처리된 모델 모두를 능가하는 성능을 보였으며, 특히 난이도가 높은 작업에서 가장 큰 향상을 나타냈습니다. 추가 분석 결과, LongCrafter 데이터는 더욱 다양하며 다양한 난이도 수준에 걸쳐 균등하게 분포되어 있으며, 학습된 모델은 위치에 관계없이 증거를 안정적으로 찾아내어 '중간 부분에서 정보 손실' 문제를 효과적으로 완화합니다.
Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision. We propose \textbf{LongCrafter}, a structured synthesis framework that couples a hierarchical task taxonomy with an evidence-grounded pipeline. The taxonomy organizes long-context understanding into local/shallow and global/deep levels and yields 32 fine-grained task types that serve as a global generative prior. Guided by this taxonomy, LongCrafter constructs task-aligned long contexts, decomposes them into explicit evidence graphs that model cross-paragraph dependencies, and generates instruction--response pairs strictly grounded in the located evidence spans, ensuring both controllable difficulty and faithful, traceable reasoning. Models fine-tuned on LongCrafter data outperform all SFT baselines and even the official post-trained models on LongBench, LongBench~v2, and LooGLE across both Qwen2.5-7B and LLaMA-3.1-8B, with the largest gains on high-difficulty tasks. Further analysis shows that LongCrafter data is more diverse and better spread across difficulty levels, and that the trained models locate evidence robustly regardless of position, effectively mitigating the ``lost in the middle'' problem.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.