심층 연구는 신뢰할 만한가? 오해를 불러일으키는 정보가 잘못된 결론으로 이어지는 현상
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
심층 연구 에이전트는 계획, 정보 검색, 증거 종합 및 보고서 생성 등 장기적인 워크플로우를 수행하는 LLM 기반 어시스턴트를 확장하지만, 개방형 정보 환경에서의 신뢰성은 아직 충분히 연구되지 않았습니다. 주요 문제는 겉으로는 신뢰할 만해 보이지만 사실상 잘못된 정보를 이러한 워크플로우에서 접하게 될 경우, 최종 보고서에서 이를 허위 결론으로 채택할 수 있는가 하는 점입니다. 이 문제점을 분석하기 위해, 심층 연구 작업에 사용되는 오해를 불러일으키는 정보를 생성하고 검증하는 프레임워크인 MisKnow-Agent를 소개합니다. MisKnow-Agent는 제어 가능한 수준의 신뢰도와 스타일을 가진 오해의 소지가 있는 정보를 생성하며, 이를 통해 DeepResearch Benchmark 작업을 기반으로 5,933개의 품질 관리된 데이터를 구축했습니다. 공개 및 비공개 심층 연구 에이전트를 대상으로 실시한 광범위한 실험 결과, 오해의 소지가 있는 정보에 노출되는 정도가 미미하더라도 최종 보고서에서 허위 결론 채택을 유발할 수 있으며, 이는 현재 심층 연구 에이전트의 전반적인 신뢰성 취약점을 드러냅니다. 검색 기능을 갖춘 검증 모델은 특정 데이터 세트에 대한 집중적 검증 과정에서 해당 정보들이 오해의 소지가 있음을 꾸준히 식별하지만, 이러한 정보들은 장기적인 연구 워크플로우에서 여전히 채택될 수 있으며, 이는 집중적 검증과 워크플로우 수준에서의 증거 활용 간의 불일치를 보여줍니다. 마지막으로, 연구 전후에 적용할 수 있는 다양한 방어 기법들을 개별적으로 그리고 조합하여 평가한 결과, 모든 구성이 허위 결론 채택을 완화하지만 완전히 막지는 못하는 것으로 나타났습니다. 이러한 결과를 바탕으로, 심층 연구의 신뢰성을 확보하기 위해서는 계획, 정보 검색, 증거 통합 또는 보고서 생성 능력 향상 외에도 모델 및 프레임워크 수준에서 증거 검증 및 수정 기능을 갖추는 것이 필수적입니다.
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, yet their reliability in open information environments remains underexplored. A key concern is whether apparently credible but factually misleading knowledge encountered in such environments can propagate through these workflows and be adopted as false conclusions in final reports. To study this failure mode, we introduce MisKnow-Agent, a framework for constructing and validating misleading knowledge for Deep Research tasks. MisKnow-Agent generates misleading instances with controllable authority levels and styles, yielding 5,933 quality-controlled instances built on DeepResearch Benchmark tasks. Extensive experiments across open-source and closed-source Deep Research agents show that even limited exposure to misleading knowledge can induce false-conclusion adoption in final reports, revealing a broad reliability vulnerability in current Deep Research agents. Although search-enabled verifier models consistently identify the retained instances as misleading during focused corpus validation, the same instances can still be adopted during long-horizon research, revealing a disconnect between focused verification and workflow-level evidence use. Finally, we evaluate pre- and post-research defenses, both individually and in combination, finding that all three configurations mitigate but do not fully prevent false-conclusion adoption. Our findings suggest that reliable Deep Research requires evidence verification and correction capabilities at both the model and framework levels, beyond improvements in planning, retrieval, evidence integration, or report-generation abilities.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.