추론의 '이유'를 연결하다: LLM에서의 추론적 추론에 대한 통일된 분류 체계 및 조사
Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs
인간의 발견과 이해에 있어 근본적인 역할을 수행함에도 불구하고, 추론적 추론—즉, 관찰에 대한 가장 그럴듯한 설명을 추론하는 과정—은 대규모 언어 모델(LLM)에서 상대적으로 덜 연구되어 왔습니다. LLM의 빠른 발전에도 불구하고, 추론적 추론과 그 다양한 측면에 대한 탐구는 아직까지 통합적인 접근보다는 단편적인 연구에 머물러 있었습니다. 본 논문은 LLM에서의 추론적 추론에 대한 최초의 종합적인 조사를 제시하며, 철학적 기초부터 현대 AI 구현에 이르기까지 그 발전 과정을 추적합니다. 본 연구는 해당 분야에서 널리 퍼져있는 개념적 혼란과 단편적인 작업 정의 문제를 해결하기 위해, 기존 연구를 체계적으로 분류하는 통일된 두 단계 정의를 제시합니다. 이 정의는 추론을 '가설 생성'—모델이 인식적 격차를 해소하여 후보 설명을 생성하는 단계—과 '가설 선택'—생성된 후보들이 평가되고 가장 그럴듯한 설명이 선택되는 단계—로 구분합니다. 이 기반을 바탕으로, 기존 연구를 추론 작업, 데이터셋, 기반 방법론, 평가 전략에 따라 분류하는 포괄적인 분류 체계를 제시합니다. 또한, 본 프레임워크의 실질적인 타당성을 검증하기 위해, 현재 LLM의 추론 작업 수행 능력을 평가하는 간결한 벤치마크 연구를 수행하고, 모델 크기, 모델 유형, 평가 방식, 그리고 생성과 선택이라는 추론 작업 유형 간의 비교 분석을 진행합니다. 나아가, 최근의 연구 결과를 종합하여, LLM의 추론적 추론 수행 능력이 연역적 및 귀납적 추론과 어떻게 관련되는지 분석함으로써, LLM의 더 넓은 추론 능력에 대한 통찰력을 제공합니다. 우리의 분석 결과는 현재 접근 방식의 중요한 한계점을 드러냅니다. 이는 정적인 벤치마크 설계, 좁은 영역 범위, 제한적인 학습 프레임워크, 그리고 추론 과정에 대한 제한적인 이해 등과 관련된 문제입니다.
Regardless of its foundational role in human discovery and sense-making, abductive reasoning--the inference of the most plausible explanation for an observation--has been relatively underexplored in Large Language Models (LLMs). Despite the rapid advancement of LLMs, the exploration of abductive reasoning and its diverse facets has thus far been disjointed rather than cohesive. This paper presents the first survey of abductive reasoning in LLMs, tracing its trajectory from philosophical foundations to contemporary AI implementations. To address the widespread conceptual confusion and disjointed task definitions prevalent in the field, we establish a unified two-stage definition that formally categorizes prior work. This definition disentangles abduction into \textit{Hypothesis Generation}, where models bridge epistemic gaps to produce candidate explanations, and \textit{Hypothesis Selection}, where the generated candidates are evaluated and the most plausible explanation is chosen. Building upon this foundation, we present a comprehensive taxonomy of the literature, categorizing prior work based on their abductive tasks, datasets, underlying methodologies, and evaluation strategies. In order to ground our framework empirically, we conduct a compact benchmark study of current LLMs on abductive tasks, together with targeted comparative analyses across model sizes, model families, evaluation styles, and the distinct generation-versus-selection task typologies. Moreover, by synthesizing recent empirical results, we examine how LLM performance on abductive reasoning relates to deductive and inductive tasks, providing insights into their broader reasoning capabilities. Our analysis reveals critical gaps in current approaches--from static benchmark design and narrow domain coverage to narrow training frameworks and limited mechanistic understanding of abductive processes...
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.