자기 회귀의 제약을 극복하는 방법: LLM을 위한 동적 인식 엔트로피 기반의 지우기 가능한 강화 학습
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs
강화 학습(RL)은 대규모 언어 모델(LLM)의 인지 능력을 확장했지만, 장기간의 논리적 추론에서 자기 회귀의 제약에 취약한 경우가 많습니다. 생성 초기에 발생하는 작은 인식 변동이 마르코프 의사 결정 프로세스를 따라 되돌릴 수 없이 전파되어 연쇄적인 오류를 유발하고 추론 경로를 붕괴로 이끌 수 있습니다. 이러한 자기 회귀 연쇄 현상을 극복하기 위해, 우리는 동적 인식 엔트로피 기반의 지우기 가능한 강화 학습($ ext{E}^3 ext{RL}$)을 제안합니다. $ ext{E}^3 ext{RL}$은 외부 신호에 대한 의존성을 없애고 모델의 고유한 로컬 자기 회귀 교차 엔트로피를 인식 불확실성의 내재적인 좌표로 활용합니다. 세그먼트 수준의 적응적 동적 임계값과 이점 할당을 도입하여, $ ext{E}^3 ext{RL}$은 모델이 과거의 중요한 키-값(KV) 캐시 스트림을 재사용하면서도 국소적인 논리적 결함을 정확하게 제거할 수 있도록 하여 추론 과정에 자체 복구 능력을 부여합니다. 우리는 $ ext{E}^3 ext{RL}$을 DeepMath-103k 데이터셋으로 학습했습니다. 실험 결과는 $ ext{E}^3 ext{RL}$이 장문 추론의 탐색 효율성을 개선하고 샘플 효율성을 향상시키면서도 선형적인 메모리 오버헤드를 유지한다는 것을 보여줍니다. AIME과 같은 수학적 추론 벤치마크에서, $ ext{E}^3 ext{RL}$은 상당한 성능 향상을 달성했으며, 특히 4B 및 8B 파라미터 모델이 기존 최고 성능(SOTA) 결과를 각각 5.349% 및 6.514% 상회하는 것으로 나타났습니다. 이러한 결과는 $ ext{E}^3 ext{RL}$이 장문 추론에서의 자기 회귀 제약을 극복하고, 차세대 자체 복구 인공 일반 지능(AGI)을 위한 이론적 및 시스템 수준의 기반을 구축한다는 것을 시사합니다.
Although reinforcement learning (RL) has expanded the cognitive boundaries of large language models (LLMs), it often remains vulnerable to the autoregressive curse in long-horizon logical reasoning: small epistemic perturbations introduced early in generation can propagate irreversibly along the Markov decision process flow, triggering cascading failures that drive the reasoning trajectory toward collapse. To overcome this autoregressive cascade, in which a single early mistake can compromise all subsequent reasoning steps, we propose dynamic epistemic entropy orchestrated erasable reinforcement learning ($\text{E}^3\text{RL}$). $\text{E}^3\text{RL}$ eliminates reliance on external signals by grounding the model's endogenous local autoregressive cross-entropy as an intrinsic coordinate of epistemic uncertainty. By introducing segment-level adaptive dynamic thresholds and advantage allocation, $\text{E}^3\text{RL}$ enables the model to precisely excise localized logical defects while reusing historical key-value (KV) cache streams, thereby endowing the reasoning process with a self-healing capability. We train $\text{E}^3\text{RL}$ on the DeepMath-103k dataset. Experimental results show that $\text{E}^3\text{RL}$ reshapes the exploration efficiency of long-sequence reasoning and improves sample efficiency while maintaining linear memory overhead. On mathematical reasoning benchmarks such as AIME, $\text{E}^3\text{RL}$ achieves substantial performance gains, with the 4B and 8B parameter models surpassing previous state-of-the-art (SOTA) results by 5.349\% and 6.514\%, respectively. These findings suggest that $\text{E}^3\text{RL}$ shatters the autoregressive curse in long-sequence reasoning and establishes a theoretical and systems-level foundation for the next generation of self-healing artificial general intelligence (AGI).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.