2607.15178v1 Jul 16, 2026 cs.CL

T^2MLR: 시간적 중간 레이어 순환을 갖춘 Transformer

T^2MLR: Transformer with Temporal Middle-Layer Recurrence

Sanjeev Arora
Sanjeev Arora
Citations: 482
h-index: 11
Xingyu Zhu
Xingyu Zhu
Princeton University
Citations: 132
h-index: 4
Ziyang Cai
Ziyang Cai
Citations: 63
h-index: 4
Yinghui He
Yinghui He
Citations: 152
h-index: 4
Yihe Dong
Yihe Dong
Citations: 16
h-index: 2

Transformer 모델의 추론 능력은 자기 회귀 디코딩 방식에 의해 제한되며, 이로 인해 풍부한 숨겨진 연산이 토큰 공간 내에서 반복적으로 압축되어 중간 추론 상태가 시간 경과에 따라 유지되기 어렵습니다. 본 논문에서는 시간적 중간 레이어 순환을 갖춘 Transformer (T2MLR)라는 잠재적인 추론 아키텍처를 제안합니다. T2MLR은 이전 토큰에서 가져온 캐시된 중간 레이어 표현을 현재 토큰 위치의 이전 레이어로 직접 통합하여, 추상적인 중간 연산이 디코딩 단계 동안 거의 인공 지능 처리 비용 없이 지속될 수 있도록 합니다. 자연어 사전 훈련 및 다중 단추 추론 미세 조정 실험에서 T2MLR은 데이터 및 파라미터 측면에서 동일한 Transformer 모델을 기반으로 하는 기존 모델보다 일관되게 뛰어난 성능을 보였습니다. 또한, 전체 레이어가 아닌 국소적인 중간 레이어 블록(네트워크의 20% 정도)에만 순환을 적용하는 것이 전체 레이어 순환보다 더 나은 결과를 보이는 경우가 많습니다. 중요한 점은 T2MLR이 처음부터 사전 훈련할 필요가 없다는 것입니다. 기존의 사전 훈련된 17억 파라미터 Transformer 모델에 순환 경로를 추가하고 짧게 미세 조정하는 것만으로도 수학적 추론 능력이 크게 향상되어 실제 적용 가능성이 높아집니다. 이러한 결과는 Transformer에서 효과적인 잠재적 추론이 이전 연구에서처럼 모든 레이어를 반복할 필요 없이, 오히려 표적화된 중간 레이어 순환을 통해 더 강력하게 나타날 수 있음을 시사합니다.

Original Abstract

Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time. We in troduce Transformers with Temporal Middle-Layer Recurrence (T2MLR), a transformers-based latent reasoning architecture that fuses a cached middle layer representation from the previous token directly into an earlier layer of the current token position, enabling abstract intermediate computation to persist across decoding steps with little inference overhead. Across natural-language pretraining and multi-hop reasoning finetuning, T2MLR consistently outperforms data- and parameter-matched Transformer base lines. Moreover, applying recurrence to only a localized middle-layer block (as little as 20% of the network) often outperforms full-layer recurrence. Im portantly, T2MLR does not require pretraining from scratch: retrofitting the recurrent pathway into an existing pretrained 1.7B Transformer and briefly finetuning substantially improves math reasoning, lowering the barrier to practical adoption. These results suggest that effective latent reasoning in Transformers does not require looping over all layers as in previous works, but can instead emerge more strongly from targeted middle-layer recurrence.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!