2606.06252v1 Jun 04, 2026 cs.AI

테스트 시간 재구성을 통한 잠재 추론의 순환 구조 완성

Closing the Loop on Latent Reasoning via Test-Time Reconstruction

Haibo Jin
Haibo Jin
Citations: 314
h-index: 5
Ye Yu
Ye Yu
Citations: 39
h-index: 4
Xiaopeng Yuan
Xiaopeng Yuan
Citations: 11
h-index: 3
Yushun Dong
Yushun Dong
Citations: 189
h-index: 8
Peng Kuang
Peng Kuang
Citations: 9
h-index: 2
Haohan Wang
Haohan Wang
Citations: 36
h-index: 3
Lijun Yu
Lijun Yu
Citations: 35
h-index: 4

최근 연구에서는 토큰 오버헤드를 줄이고 이산적인 통신 병목 현상을 피하기 위해 자연어 추론 과정을 잠재 또는 캐시 레벨 표현으로 변환하는 경향이 있습니다. 하지만 이러한 변화는 텍스트 기반 추론의 중요한 장점을 없애는데, 바로 중간 상태를 검사할 수 없게 되어 원래 질문의 제약 조건을 잠재 상태가 여전히 유지하고 있는지 판단하기 어렵다는 것입니다. 결과적으로, 잠재 추론은 일반적으로 폐쇄 루프 시스템이 아닌 개방 루프 시스템으로 작동하며, 잠재 상태가 생성되고 소비될 때 입력에 기반한 신뢰성 검사가 이루어지지 않습니다. 본 논문에서는 ReLAT(Reconstruction-Guided Latent Reasoning At Test Time)이라는 자체 지도 학습 방식의 테스트 시간 훈련 방법을 제안합니다. 이 방법은 질문 자체를 기준으로 하여 폐쇄 루프 시스템을 구축합니다. 핵심적인 관찰 사항은, 잠재 상태가 질문을 충실하게 표현한다면, 그 질문이 잠재 상태로부터 복구 가능해야 한다는 것입니다. 반대로, 질문을 복구할 수 없다면, 잠재 상태는 작업과 관련된 정보를 잃어버린 것입니다. ReLAT는 이 원칙을 활용하여 미분 가능한 질문 -> 잠재적 사고 -> 질문의 순환 구조를 구축하고, 답변 생성 전에 잠재적 사고를 통해 질문 재구성 손실을 최적화합니다. 이를 통해 불투명한 잠재 계산을 그것이 표현해야 할 문제 사양에 연결합니다. Qwen 패밀리를 기반으로 한 수학적 추론, 지식 질의응답 및 코드 생성 벤치마크에서 ReLAT는 단일 모델 추론, 텍스트 기반 협업, 개방 루프 잠재 협업 및 대체 테스트 시간 훈련 목표보다 일관되게 성능 향상을 보였습니다. Qwen3-8B 모델에서 ReLAT는 AIME 2024 정확도를 56.7%에서 73.3%로 향상시켜, 가장 강력한 개방 루프 잠재 기준 모델 대비 16.6%p의 성능 향상을 달성했습니다.

Original Abstract

Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a discrete communication bottleneck. However, this shift also removes a key advantage of textual reasoning: intermediate states are no longer inspectable, making it difficult to determine whether a latent state still preserves the constraints of the original query. As a result, latent reasoning typically operates in an open loop, where a latent state is produced and consumed without an input-anchored fidelity check. We propose ReLAT (Reconstruction-Guided Latent Reasoning At Test Time), a self-supervised test-time training method that closes this loop using the query itself as the reference. Our key observation is that if a latent state faithfully represents a query, the query should be recoverable from it; if the query cannot be recovered, the latent state has lost task-relevant information. ReLAT operationalizes this principle by constructing a differentiable Question -> Latent Thought -> Question cycle and optimizing query reconstruction loss through the latent thought before answer generation. This anchors opaque latent computation to the problem specification it is supposed to represent. Across mathematical reasoning, knowledge QA, and code generation benchmarks on the Qwen family, ReLAT consistently improves over single-model inference, text-based collaboration, open-loop latent collaboration, and alternative test-time training objectives. On Qwen3-8B, ReLAT raises AIME 2024 accuracy from 56.7% to 73.3%, a 16.6-point gain over the strongest open-loop latent baseline.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!