2607.25915v1 Jul 28, 2026 cs.AI

페넬로페: 효율적인 구조적 추론을 위한 지역화된 잠재적 반복

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Yutong Chen
Yutong Chen
Citations: 21
h-index: 3
Shouqian Shi
Shouqian Shi
Citations: 635
h-index: 12
Zirui Ding
Zirui Ding
Citations: 0
h-index: 0
Xinran Liu
Xinran Liu
Citations: 0
h-index: 0
Haocheng Wang
Haocheng Wang
Citations: 0
h-index: 0
Jiaying Wang
Jiaying Wang
Citations: 0
h-index: 0
Tianxing Xu
Tianxing Xu
Citations: 0
h-index: 0
Yuanxi Wang
Yuanxi Wang
Citations: 0
h-index: 0

복잡한 구조적 추론 작업은 종종 추가 계산을 요구하지만, 현재의 언어 모델은 주로 파라미터 규모를 늘리거나 중간 단계를 Chain-of-Thought (CoT) 토큰으로 직렬화하여 이를 수행합니다. 전자는 훈련 및 배포 비용을 증가시키고, 후자는 추론 계산을 자기 회귀 출력 길이에 연결시킵니다. 본 연구에서는 사전 학습된 디코더 전용 Transformer 모델에 대한 효율적인 잠재적 추론 프레임워크인 페넬로페를 소개합니다. 페넬로페는 반복 계산을 선택된 디코더 구간으로 제한합니다. 하위 디코더 접두부는 한 번 평가되어 문제 조건에 맞는 경계 메모리를 구성하고, 이 메모리는 답변 생성을 위해 시간 조절 GRU 동역학과 반복 읽기 상태를 통해 반복적으로 개선됩니다. 점진적인 CoT-to-latent 교육 과정을 통해 명시적인 추론을 이 내부 반복 경로로 이전하여 추가 계산을 잠재 공간에 할당하면서 전체 디코더를 반복 실행하거나 긴 중간 추적을 생성하는 것을 방지합니다. 오픈 소스 구조적 추론 벤치마크에서 수행된 실험 결과, 페넬로페는 검증 데이터셋에서 선택한 잠재적 예산 하에서 기존의 잠재적 추론 모델과 경쟁력 있는 정확도를 달성하면서 측정된 추론 지연 시간을 줄입니다. 이러한 결과는 잠재적 개선이 좁은 디코더 구간으로 제한될 수 있으며, 긴 명시적 추론 경로를 생성하지 않고 전체 디코더 실행을 반복하는 것을 방지하여 디코더 전용 Transformer 모델에 대한 실질적인 정확도-효율성 균형을 제공한다는 것을 보여줍니다.

Original Abstract

Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-conditioned boundary memory, which is then iteratively refined through time-modulated GRU dynamics and recurrent readout states before answer generation. A progressive CoT-to-latent curriculum transfers visible reasoning into this internal recurrent path, allowing additional computation to be allocated in latent space without repeatedly executing the complete decoder or generating a long intermediate trace. Experiments on open-source structured-reasoning benchmarks show that, at validation-selected latent budgets, Penelope attains competitive accuracy relative to established latent-reasoning models while reducing measured inference latency. These results show that latent refinement can be localized to a narrow decoder interval, reducing repeated full-decoder execution without generating a long visible reasoning trace and providing a practical accuracy-efficiency tradeoff for decoder-only Transformer models.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!