DeepSurvey: 자동 설문 생성 시 분석 깊이와 인용 신뢰성 향상
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation
과학 문헌의 급격한 증가로 인해, AI 연구자와 인간 연구자 모두에게 자동 설문 생성은 핵심적인 기능이 되었습니다. 그러나 기존 시스템은 초록 및 개별 논문 처리에 의존하여 분석 깊이가 제한적이며, 부정확한 검색과 사후 검증으로 인한 신뢰할 수 없는 인용 정보로 인해 피상적인 설문을 생성하고 연구자를 오도할 수 있습니다. 본 연구에서는 이러한 문제를 해결하는 DeepSurvey라는 에이전트 기반 시스템을 제시합니다. DeepSurvey는 분석 깊이를 향상시키기 위해, 전체 텍스트 논문에서 구조화된 핵심 정보를 추출하고, 클러스터링 및 비교 분석을 통해 논문 간의 관계를 모델링하며, 코드 저장소 분석을 통합하여 구현 수준의 세부 정보를 복원합니다. 또한, 인용 신뢰성을 강화하기 위해, 주제 중심 검색을 위한 시타이션 그래프 확장과 하이브리드 필터링을 결합하고, 증거 기반의 인용 할당을 적용하며, 다중 수준 에이전트 정제 방식을 사용하여 인용 주장의 일관성을 검증합니다. 실험 결과, DeepSurvey는 가장 높은 콘텐츠 점수(8.644/10)와 인용 품질(가장 강력한 기준 모델 대비 12.3% 및 9.3%의 재현율 및 정밀도 향상)을 달성했으며, 다양한 분야에서 더 안정적으로 일반화됩니다(CS-to-non-CS 점수 감소: 0.14 vs 0.22 to 0.69). 또한, 전문가들은 DeepSurvey가 인간이 작성한 설문보다 더 우수하다고 평가했습니다(전체 품질 83.3%, 콘텐츠 깊이 100%).
As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existing systems suffer from limited analytical depth due to reliance on abstracts and isolated paper processing, and unreliable citations from imprecise retrieval and post-hoc grounding, producing superficial surveys and may mislead researchers. We present DeepSurvey, an agentic system that addresses both. To enhance depth, DeepSurvey extracts structured keynotes from full-text papers, models cross-paper relationships through clustering and comparative analysis, and integrates code-repository analysis to recover implementation-level details. To fortify reliability, it combines citation-graph expansion with hybrid filtering for topic-focussed retrieval, enforces evidence-constrained citation assignment, and deploys multi-granularity agentic refinement to validate citation-claim alignment. Experiments show that DeepSurvey achieves the highest content score (8.644/10) and citation quality (12.3% and 9.3% recall and precision gains over the strongest baseline), generalizes more robustly across domains (0.14 vs 0.22 to 0.69 CS-to-non-CS drop), and is preferred over human-written surveys by domain experts (83.3% overall quality, 100% content depth).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.