2604.16694v1 Apr 17, 2026 cs.AI

RankGuide: 텐서 랭크 기반 라우팅 및 제어를 통한 효율적인 추론

RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning

Yupeng Su
Yupeng Su
Citations: 149
h-index: 4
Zheng Zhang
Zheng Zhang
Citations: 59
h-index: 4
Souvik Kundu
Souvik Kundu
Citations: 20
h-index: 3
Jiayi Tian
Jiayi Tian
Citations: 15
h-index: 2
Ryan Solgi
Ryan Solgi
Citations: 132
h-index: 4

대규모 추론 모델(LRM)은 명시적인 다단계 사고 과정(CoT)을 생성하여 문제 해결 능력을 향상시키지만, 상당한 추론 지연 시간과 계산 오버헤드를 발생시킵니다. 이러한 문제를 완화하기 위해, 최근 연구에서는 소규모 추론 모델(SRM)이 중간 추론 단계를 생성하여 더 나은 정확도-지연 시간 균형을 달성하는 모델 협업 패러다임을 탐구했습니다. 하지만 최근의 발전에도 불구하고, 협업 시스템에서 SRM의 오류를 효과적이고 효율적으로 감지하고 완화하는 것은 여전히 중요한 과제입니다. 이 문제를 해결하기 위해, 우리는 생성된 텍스트와 숨겨진 상태 공간 모두에서 SRM 추론을 분석하고, 과신(overconfidence), 불확실성(uncertainty), 과도한 재검증(heavy revalidation)의 세 가지 유형의 오류 모드를 식별했습니다. 이러한 통찰력을 바탕으로, 우리는 텐서 랭크 기반 라우팅 및 제어를 통해 SRM-LRM 협업의 효율성과 효과를 향상시키는 프레임워크인 **RankGuide**를 제안합니다. 구체적으로, RankGuide는 연속적인 숨겨진 상태에서 파생된 텐서 랭크 신호를 포함하는 라우팅 신호를 활용하여 SRM이 오류를 일으킬 가능성이 높을 때 감지하고, 선택적으로 LRM을 호출합니다. 또한, 텐서 랭크 필터링된 제어 벡터 추출 방법을 도입하여 SRM의 추론 경로를 조절함으로써 생성 품질을 향상시킵니다. RankGuide는 텐서 랭크 신호를 활용하여 라우팅과 제어를 모두 개선함으로써, SRM-LRM 협업 시스템이 더 적은 단계로 더 효율적인 추론을 수행하고 향상된 정확도를 달성할 수 있도록 합니다. 여러 추론 벤치마크에 대한 실험 결과는 RankGuide가 LRM에 비해 최대 1.75배까지 지연 시간을 줄이고, 이전 방법과 경쟁력 있는 정확도를 유지하는 것을 보여줍니다.

Original Abstract

Large reasoning models (LRMs) enhance problem-solving capabilities by generating explicit multi-step chains of thought (CoT) reasoning; however, they incur substantial inference latency and computational overhead. To mitigate this issue, recent works have explored model collaboration paradigms, where small reasoning models (SRMs) generate intermediate reasoning steps to achieve a better accuracy--latency trade-off. Despite recent progress, effectively and efficiently detecting and mitigating SRM failures in collaborative systems remains a key challenge. To address this issue, we analyze SRM inference in both the generated text and hidden-state spaces, and identify three types of failure modes: \textit{overconfidence}, \textit{uncertainty}, and \textit{heavy revalidation}. Building on these insights, we propose \textbf{RankGuide}, a framework that improves the efficiency and effectiveness of SRM--LRM collaboration through tensor-rank-guided routing and steering. Specifically, RankGuide leverages a routing signal that incorporates tensor-rank signals derived from consecutive hidden states to detect when SRMs are likely to fail and selectively invoke LRMs. In addition, we introduce a tensor-rank-filtered steering vector extraction method to modulate the reasoning trajectory of SRMs, thereby improving their generation quality. By improving both routing and steering through tensor-rank signals, RankGuide enables SRM--LRM collaborative systems to achieve more efficient reasoning with fewer steps and improved accuracy. Experiments across three reasoning domains -- mathematics, code generation, and scientific QA -- demonstrate the efficacy of RankGuide in reducing latency by up to $1.75\times$ compared to LRM, while maintaining competitive accuracy relative to prior methods. The code is available at \href{https://github.com/TTTTTTris/RankGuide}{https://github.com/TTTTTTris/RankGuide}.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!