2604.25665v1 Apr 28, 2026 cs.CL

LLM-ReSum: 자기 평가를 통한 LLM 기반 요약 생성 프레임워크

LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation

Haihua Chen
Haihua Chen
Citations: 143
h-index: 7
Huyen Nguyen
Huyen Nguyen
Citations: 14
h-index: 2
Haoxuan Zhang
Haoxuan Zhang
Citations: 90
h-index: 6
Yang Zhang
Yang Zhang
Citations: 51
h-index: 3
Junhua Ding
Junhua Ding
Citations: 138
h-index: 7

대규모 언어 모델(LLM)이 생성한 요약의 신뢰성 있는 평가는 여전히 해결해야 할 과제이며, 특히 다양한 도메인과 문서 길이에 걸쳐 더욱 어렵습니다. 본 연구에서는 5개 도메인에 걸쳐 7개의 데이터 세트를 사용하여 14개의 자동 요약 평가 지표 및 LLM 기반 평가기를 종합적으로 평가했습니다. 데이터 세트는 짧은 뉴스 기사부터 긴 과학, 정부, 법률 문서(2K-27K 단어)까지 포함하며, 1,500개 이상의 사람이 직접 작성한 요약이 포함되어 있습니다. 연구 결과, 전통적인 어휘 중복 기반 평가 지표(예: ROUGE, BLEU)는 사람의 판단과 약하거나 부정적인 상관관계를 보이는 반면, 특정 작업에 특화된 신경망 기반 평가 지표 및 LLM 기반 평가기는 언어적 품질 평가 측면에서 훨씬 더 높은 일관성을 보입니다. 이러한 결과를 바탕으로, 본 연구에서는 모델 미세 조정 없이 LLM 기반 평가와 생성을 통합하여 폐쇄 루프 피드백 시스템을 구축하는 자기 성찰 요약 생성 프레임워크인 LLM-ReSum을 제안합니다. 세 가지 도메인에서 LLM-ReSum은 사실 정확도를 최대 33%, 보편성을 최대 39% 향상시켜 저품질 요약을 개선했습니다. 또한, 인간 평가자는 89%의 경우 개선된 요약을 선호했습니다. 또한, 법률 문서 요약에 대한 새로운 인간 주석 벤치마크인 PatentSumEval을 새롭게 제시하며, 이는 180개의 전문가가 평가한 요약으로 구성되어 있습니다. 모든 코드 및 데이터 세트는 GitHub에서 공개될 예정입니다.

Original Abstract

Reliable evaluation of large language model (LLM)-generated summaries remains an open challenge, particularly across heterogeneous domains and document lengths. We conduct a comprehensive meta-evaluation of 14 automatic summarization metrics and LLM-based evaluators across seven datasets spanning five domains, covering documents from short news articles to long scientific, governmental, and legal texts (2K-27K words) with over 1,500 human-annotated summaries. Our results show that traditional lexical overlap metrics (e.g., ROUGE, BLEU) exhibit weak or negative correlation with human judgments, while task-specific neural metrics and LLM-based evaluators achieve substantially higher alignment, especially for linguistic quality assessment. Leveraging these findings, we propose LLM-ReSum, a self-reflective summarization framework that integrates LLM-based evaluation and generation in a closed feedback loop without model finetuning. Across three domains, LLM-ReSum improves low-quality summaries by up to 33% in factual accuracy and 39% in coverage, with human evaluators preferring refined summaries in 89% of cases. We additionally introduce PatentSumEval, a new human-annotated benchmark for legal document summarization comprising 180 expert-evaluated summaries. All code and datasets will be released in GitHub.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!