Think$^{2}$: 대규모 언어 모델에서의 기반형 메타인지 추론
Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models
대규모 언어 모델(LLM)은 강력한 추론 성능을 보여주지만, 스스로의 오류를 안정적으로 모니터링하고 진단하며 수정하는 능력은 여전히 제한적이다. 본 연구에서는 앤 브라운(Ann Brown)의 조절 주기(계획, 모니터링, 평가)를 구조화된 프롬프팅 아키텍처로 구현한 심리학적 기반의 메타인지 프레임워크를 도입하고, 이를 적응적 노력 할당을 위한 경량화된 이중 과정 메타컨트롤러(MetaController)에 통합하는 방법을 연구한다. Llama-3 및 Qwen-3(8B)를 활용해 다양한 추론 및 진단 벤치마크(GSM8K, CRUXEval, MBPP, AIME, CorrectBench, TruthfulQA)를 테스트한 결과, 명시적인 조절 구조화가 오류 진단 능력을 실질적으로 향상시키고 성공적인 자가 수정 비율을 3배 증가시키는 것으로 나타났다. 580개의 질의 쌍에 대한 블라인드 인간 평가에서는 신뢰성 및 메타인지적 자아 인식 측면에서 표준 및 생각의 사슬(Chain-of-Thought) 베이스라인 대비 84%의 종합적인 선호도를 보였다. LLM의 추론을 확립된 인지 이론에 근거하도록 하는 것은 보다 투명하고 진단 측면에서 견고한 AI 시스템을 향한 원칙적인 방향을 제시한다.
Large Language Models (LLMs) demonstrate strong reasoning performance, yet their ability to reliably monitor, diagnose, and correct their own errors remains limited. We introduce a psychologically grounded metacognitive framework that operationalizes Ann Brown's regulatory cycle (Planning, Monitoring, and Evaluation) as a structured prompting architecture, and study its integration within a lightweight dual-process MetaController for adaptive effort allocation. Across diverse reasoning and diagnostic benchmarks (GSM8K, CRUXEval, MBPP, AIME, CorrectBench, and TruthfulQA) using Llama-3 and Qwen-3 (8B), explicit regulatory structuring substantially improves error diagnosis and yields a threefold increase in successful self-correction. Blinded human evaluations over 580 query pairs show an 84% aggregate preference for trustworthiness and metacognitive self-awareness over standard and Chain-of-Thought baselines. Grounding LLM reasoning in established cognitive theory offers a principled path toward more transparent and diagnostically robust AI systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.