혼합 및 매칭: 확장 가능한 주제 기반 교육 요약화를 위한 문맥 쌍 사용
Mix and Match: Context Pairing for Scalable Topic-Controlled Educational Summarisation
주제 기반 요약은 사용자가 소스 문서의 특정 측면에 초점을 맞춘 요약을 생성할 수 있도록 합니다. 본 논문에서는 주제 기반 요약화를 수행하기 위한 작은 언어 모델(sLM)을 학습시키는 데이터 증강 전략을 연구합니다. 우리는 다양한 문서의 문맥을 결합하여 대비 학습 예제를 생성하는 쌍 기반 데이터 증강 방법을 제안합니다. 이를 통해 모델이 주제와 요약 간의 관계를 보다 효과적으로 학습할 수 있습니다. Wikipedia에서 파생된 주제로 확장된 SciTLDR 데이터 세트를 사용하여, 증강 규모가 모델 성능에 미치는 영향을 체계적으로 평가했습니다. 결과는 증강 규모가 증가함에 따라 윈율과 의미적 정렬성이 지속적으로 향상되는 것을 보여줍니다. 동시에 실제 학습 데이터의 양은 일정하게 유지됩니다. 결과적으로, 제안된 증강 방법을 사용하여 학습된 T5-base 모델은 더 큰 모델과 경쟁력 있는 성능을 달성합니다. 이는 훨씬 적은 파라미터와 실제 학습 예제를 사용했음에도 불구하고 얻어진 결과입니다.
Topic-controlled summarisation enables users to generate summaries focused on specific aspects of source documents. This paper investigates a data augmentation strategy for training small language models (sLMs) to perform topic-controlled summarisation. We propose a pairwise data augmentation method that combines contexts from different documents to create contrastive training examples, enabling models to learn the relationship between topics and summaries more effectively. Using the SciTLDR dataset enriched with Wikipedia-derived topics, we systematically evaluate how augmentation scale affects model performance. Results show consistent improvements in win rate and semantic alignment as the augmentation scale increases, while the amount of real training data remains fixed. Consequently, a T5-base model trained with our augmentation approach achieves competitive performance relative to larger models, despite using significantly fewer parameters and substantially fewer real training examples.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.