2606.32032v1 Jun 30, 2026 cs.CL

메타인지 피드백을 활용한 강화 학습은 LLM에서 신뢰할 수 있는 불확실성 표현을 이끌어낸다

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

G. Yona
G. Yona
Citations: 4,440
h-index: 17
Avi Caciularu
Avi Caciularu
Citations: 3,483
h-index: 9
Idan Szpektor
Idan Szpektor
Citations: 9,722
h-index: 37
Arman Cohan
Arman Cohan
Citations: 32
h-index: 3
Gabrielle Kaili-May Liu
Gabrielle Kaili-May Liu
Yale University
Citations: 100
h-index: 4

메타 인지(Metacognition)는 자신의 인지 과정을 감시하고 조절하는 능력으로, 지능의 중요한 구성 요소입니다. 하지만 현재의 LLM은 메타인지 역량에서 심각한 결함을 보입니다. 즉, 높은 확신을 가지고 환상을 생성하고, 지식 경계를 인식하지 못하며, 내부적인 불확실성을 잘못 표현하여 신뢰성과 안정성을 저해합니다. 과제 수행 결과를 모니터링하고 그에 따라 행동을 조정하는 것은 메타 인지의 핵심이기 때문에, 자신의 성능을 정확하게 판단할 수 있는 모델이 성능 향상에 더 유리하다고 가정했습니다. 우리는 이 아이디어를 두 가지 새로운 메커니즘을 통해 구현했습니다: 첫째, 강화 학습과 메타인지 피드백(RLMF)은 모델의 자체 평가 품질을 기반으로 선호도 최적화 과정에서 완성 순위를 개선하는 방법입니다. 둘째, 메타인지 데이터 선택은 유사한 자체 평가를 사용하여 고가치 훈련 예제를 식별하며, 이는 일반적인 능동 학습보다 뛰어난 성능을 보입니다. 우리는 이러한 혁신을 '진실된 교정(faithful calibration)'이라는 메타인지적 과제에 적용했습니다. 이 과제의 목표는 표현되는 불확실성과 내재적인 불확실성을 일치시키는 것이며, 이는 최첨단 LLM에게도 어려운 문제입니다. 우리는 두 단계로 분리된 접근 방식을 채택하여, 먼저 RLMF를 사용하여 모델이 보고하는 신뢰도 점수의 진실성을 교정하고, 그 다음 특정 출력 편집을 통해 자연스럽고 상황에 맞는 언어적 불확실성을 표현합니다. 광범위한 실험 결과, RLMF는 다양한 작업에서 일반화 가능하고 최첨단 수준의 진실된 교정을 달성하면서 정확도를 유지하는 것으로 나타났습니다. 또한, RLMF는 표준 강화 학습보다 최대 63% 더 우수한 성능을 보이며, 모델이 자신의 능력 한계를 평가하고 표현하는 능력을 향상시킵니다. 이러한 결과는 RLMF가 LLM의 메타 인지 능력을 향상시켜 더욱 발전된 기능과 일관성을 제공하는 유망한 방법론임을 시사하며, 또한 메타인지적 성능을 효과적인 강화 학습 신호로 활용하여 기존의 내재적 피드백 방법의 한계를 극복할 수 있음을 제안합니다.

Original Abstract

Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes. Yet LLMs exhibit systemic deficiencies in key metacognitive faculties: they hallucinate with high confidence, fail to recognize knowledge boundaries, and misrepresent their internal uncertainty--undermining trustworthiness and reliability. Since monitoring task performance and adapting behavior accordingly are central to metacognition, we posit that models capable of accurately judging their own performance are better positioned to improve it. We operationalize this idea via two novel mechanisms: reinforcement learning with metacognitive feedback (RLMF), a paradigm to refine completion rankings during preference optimization based on the quality of a model's self-judgments of performance, and metacognitive data selection, which uses similar self-judgments to identify high-value training examples, outperforming naive active learning. We apply these innovations to the problem of faithful calibration (FC), a task that is itself fundamentally metacognitive: the goal is to align expressed with intrinsic uncertainty, difficult even for frontier LLMs. We adopt a two-stage, decoupled approach, first using these methods to calibrate the faithfulness of models' self-reported confidence scores, then mapping to natural, context-adaptable linguistic uncertainty via targeted output editing. Extensive experiments show RLMF achieves generalizable, state-of-the-art FC on diverse tasks while preserving accuracy. Further, RLMF surpasses standard RL by up to 63% while enhancing models' ability to assess and express their own capability limits. This positions RLMF as a promising paradigm to enhance LLM metacognition toward improved abilities and alignment, and suggests metacognitive performance as an effective RL signal to overcome limits of prior intrinsic feedback methods.

1 Citations
0 Influential
18.5 Altmetric
93.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!