2605.29522v1 May 28, 2026 cs.AI

DeepSurvey: 자동 설문 생성 시 분석 깊이와 인용 신뢰성 향상

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation

Da Ma
Da Ma
Citations: 469
h-index: 10
Lu Chen
Lu Chen
Citations: 460
h-index: 12
Kai Yu
Kai Yu
Citations: 542
h-index: 13
Hanqi Li
Hanqi Li
Citations: 44
h-index: 4
Yunzhe Zhang
Yunzhe Zhang
Citations: 41
h-index: 2
Ziyue Yang
Ziyue Yang
Citations: 23
h-index: 2
Zijian Hu
Zijian Hu
Scale AI
Citations: 345
h-index: 6
Tiancheng Huang
Tiancheng Huang
Citations: 50
h-index: 4
Chenrun Wang
Chenrun Wang
Citations: 113
h-index: 1
Xiaobao Wu
Xiaobao Wu
Citations: 189
h-index: 7
Zijian Wang
Zijian Wang
Citations: 4,581
h-index: 12

과학 문헌의 급격한 증가로 인해, AI 연구자와 인간 연구자 모두에게 자동 설문 생성은 핵심적인 기능이 되었습니다. 그러나 기존 시스템은 초록 및 개별 논문 처리에 의존하여 분석 깊이가 제한적이며, 부정확한 검색과 사후 검증으로 인한 신뢰할 수 없는 인용 정보로 인해 피상적인 설문을 생성하고 연구자를 오도할 수 있습니다. 본 연구에서는 이러한 문제를 해결하는 DeepSurvey라는 에이전트 기반 시스템을 제시합니다. DeepSurvey는 분석 깊이를 향상시키기 위해, 전체 텍스트 논문에서 구조화된 핵심 정보를 추출하고, 클러스터링 및 비교 분석을 통해 논문 간의 관계를 모델링하며, 코드 저장소 분석을 통합하여 구현 수준의 세부 정보를 복원합니다. 또한, 인용 신뢰성을 강화하기 위해, 주제 중심 검색을 위한 시타이션 그래프 확장과 하이브리드 필터링을 결합하고, 증거 기반의 인용 할당을 적용하며, 다중 수준 에이전트 정제 방식을 사용하여 인용 주장의 일관성을 검증합니다. 실험 결과, DeepSurvey는 가장 높은 콘텐츠 점수(8.644/10)와 인용 품질(가장 강력한 기준 모델 대비 12.3% 및 9.3%의 재현율 및 정밀도 향상)을 달성했으며, 다양한 분야에서 더 안정적으로 일반화됩니다(CS-to-non-CS 점수 감소: 0.14 vs 0.22 to 0.69). 또한, 전문가들은 DeepSurvey가 인간이 작성한 설문보다 더 우수하다고 평가했습니다(전체 품질 83.3%, 콘텐츠 깊이 100%).

Original Abstract

As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existing systems suffer from limited analytical depth due to reliance on abstracts and isolated paper processing, and unreliable citations from imprecise retrieval and post-hoc grounding, producing superficial surveys and may mislead researchers. We present DeepSurvey, an agentic system that addresses both. To enhance depth, DeepSurvey extracts structured keynotes from full-text papers, models cross-paper relationships through clustering and comparative analysis, and integrates code-repository analysis to recover implementation-level details. To fortify reliability, it combines citation-graph expansion with hybrid filtering for topic-focussed retrieval, enforces evidence-constrained citation assignment, and deploys multi-granularity agentic refinement to validate citation-claim alignment. Experiments show that DeepSurvey achieves the highest content score (8.644/10) and citation quality (12.3% and 9.3% recall and precision gains over the strongest baseline), generalizes more robustly across domains (0.14 vs 0.22 to 0.69 CS-to-non-CS drop), and is preferred over human-written surveys by domain experts (83.3% overall quality, 100% content depth).

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!