OneLatent: 시각적 잠재 추론을 위한 단일 토큰 압축
OneLatent: Single-Token Compression for Visual Latent Reasoning
CoT(Chain-of-thought) 프롬프팅은 추론 능력을 향상시키지만, 종종 추론 비용을 한두 자릿수(orders of magnitude)만큼 증가시킵니다. 이러한 문제를 해결하기 위해, 우리는 렌더링된 CoT 이미지와 DeepSeek-OCR 은닉 상태(hidden states)를 이용한 지도를 통해 중간 추론 과정을 단 하나의 잠재 토큰으로 압축하는 프레임워크인 OneLatent를 제안합니다. 텍스트 단계를 이미지로 렌더링함으로써, 모델이 장황한 텍스트 근거를 출력할 필요 없이 검사 및 감사가 가능한 결정론적 지도 신호를 얻을 수 있습니다. 여러 벤치마크에서 OneLatent는 텍스트 CoT 대비 평균 정확도 하락은 2.21%에 불과하면서 평균 출력 길이를 11배 줄였고, 출력 토큰 기여도(OTC)는 6.8배 향상시켰습니다. 긴 연쇄 논리 추론의 경우, OneLatent는 단 하나의 잠재 토큰으로 ProntoQA에서 99.80%, ProsQA에서 97.80%를 달성했으며, 최대 87.4배의 압축률을 보이며 압축 제약 일반화를 지원합니다.
Chain-of-thought (CoT) prompting improves reasoning but often increases inference cost by one to two orders of magnitude. To address these challenges, we present \textbf{OneLatent}, a framework that compresses intermediate reasoning into a single latent token via supervision from rendered CoT images and DeepSeek-OCR hidden states. By rendering textual steps into images, we obtain a deterministic supervision signal that can be inspected and audited without requiring the model to output verbose textual rationales. Across benchmarks, OneLatent reduces average output length by $11\times$ with only a $2.21\%$ average accuracy drop relative to textual CoT, while improving output token contribution (OTC) by $6.8\times$. On long-chain logical reasoning, OneLatent reaches $99.80\%$ on ProntoQA and $97.80\%$ on ProsQA with one latent token, with compression up to $87.4\times$, supporting compression-constrained generalization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.